
d-Matrix Raptor XPUs Join Nvidia's NVLink Fusion Racks
d-Matrix will build its Raptor inference XPUs around Nvidia NVLink Fusion and MGX racks, with 3 TB/s per-XPU bandwidth and availability in Q4 2027.
What d-Matrix Announced
Inference chip startup d-Matrix said on September 10 that its next-generation Raptor XPUs will be built around Nvidia's NVLink Fusion interconnect and the MGX rack-scale reference designs. The agreement covers a multi-year product roadmap and puts d-Matrix silicon inside the same racks, switch fabrics and cooling that Nvidia's own AI infrastructure uses.
- Raptor XPUs are designed from the ground up for NVLink Fusion and the Nvidia MGX rack ecosystem
- 3 TB/s of all-to-all bandwidth per XPU over sixth-generation NVLink, according to Nvidia
- Nvidia cites 3x lower XPU-to-XPU latency and 10x higher packet rates than standard Ethernet for the scale-up domain
- Initial availability is expected in Q4 2027, with Raptor taping out before the end of this year
The pairing extends past the interconnect. d-Matrix plans to pair its accelerators with Nvidia Vera CPUs, BlueField and ConnectX network cards, and SpectrumX Ethernet, according to The Register's report on the announcement.
Why an AI Chip Startup Would Adopt a Competitor's Rack
The obvious reading is that a challenger just plugged into the incumbent's ecosystem. The more useful reading is about what a startup gets to stop building. Rack-scale AI infrastructure is not only silicon — it is the power delivery, the liquid cooling loop, the switch fabric, the mechanical standard that a data centre operator has already qualified and stocked spares for. Designing around MGX means d-Matrix customers deploy into racks and NVSwitch fabrics they already run.
Sid Sheth, d-Matrix cofounder and CEO, framed it in terms of constraints: "Demand for inference is soaring, but capital, time and energy remain finite." His stated goal is a "faster, lower-risk path to deploy and scale ultralow-latency inference" — which is a candid description of what interoperability buys a company that is not Nvidia.
It is also part of a pattern. NVLink Fusion partners now include AWS, Arm, Intel, Fujitsu and Samsung, and Nvidia has been building out the ecosystem financially as well — we covered its $3.5 billion MediaTek investment for custom AI chips last month.
What Makes the Raptor XPU Different?
d-Matrix's technical bet is in-memory compute: 3D-stacked DRAM bonded directly on top of the compute logic, trading capacity for bandwidth that sits much closer to SRAM speeds than to conventional DRAM. The Register reports each Raptor card carries 32GB of 3D-stacked DRAM delivering roughly 100 TB/s of memory bandwidth, with the current-generation XPUs in d-Matrix's NVL144 racks holding about 16GB and roughly 50 TB/s.
At rack scale, The Register's figures put 144 Raptor accelerators at approximately 2.3TB of total memory capacity and 7.2 petabytes per second of peak aggregate memory bandwidth — enough, on d-Matrix's arithmetic, for models above four trillion parameters at 4-bit precision. Those are vendor and reporter figures for hardware that has not taped out yet, so the sensible posture is interest rather than conclusions.
The target workloads are the latency-sensitive ones: coding assistants, chatbots and voice agents, where the time to first token is what users actually feel. That is a different design point from raw training throughput, and it is the same direction custom inference silicon has been moving all year — see Meta's MTIA 400 and its FP4 inference push. More hardware coverage in our AI section.
Sources: NVIDIA Blog — September 10, 2026; The Register — September 10, 2026.
More AI Stories

Gemini 3.8 Live Adds Reasoning to Real-Time Voice AI
Google's new Gemini 3.8 Live models reason mid-conversation across 97 languages and top the speech quality index at 82.6. Here is what changes.

Perplexity Portable Computer Runs Local AI on Windows
Perplexity's on-device agent now runs on Windows PCs with 24GB+ RTX GPUs, keeping the model, harness, orchestrator and scheduler off the cloud.

Google AI Hits 300 Languages, Covering 86% of People
Google says its products now work in 300+ languages for 7 billion people, backed by open speech datasets covering 109 Indian and 27 African languages.
