As an NVIDIA NVLink Fusion partner, d-Matrix will work closely with NVIDIA to incorporate d-Matrix's next-gen inference XPUs, starting with d-Matrix Raptor™, directly into the latest NVIDIA rack reference architecture design featuring NVIDIA Vera CPUs, NVIDIA NVLink switches, NVIDIA BlueField-4 DPUs, NVIDIA ConnectX-9 SuperNICs, and NVIDIA Spectrum-X Ethernet networking. d-Matrix is also partnering with Astera Labs, a connectivity solution leader within the NVIDIA NVLink Fusion ecosystem, to deliver custom solutions to ensure high-throughput, seamless data flow throughout the system. The d-Matrix rack, enabled by the MGX platform, will feature modular cable-free trays built with NVIDIA's mature, proven MGX ecosystem and supply chain for fast and seamless deployments.

NVLink Fusion gives d-Matrix a mature, high-bandwidth, low-latency scale-up foundation for connecting d-Matrix XPUs to NVIDIA rack-scale infrastructure. With NVLink Fusion and the MGX ecosystem, d-Matrix can build around the same rack architecture, networking and supply-chain used across the NVIDIA platform. The first engagement point in the collaboration will be d-Matrix Raptor XPUs plugging into the NVIDIA MGX rack resulting in higher performance, greater deployment flexibility for customers, and a scalable architecture for expanding Raptor-based inference clusters as demand grows.

"This collaboration with NVIDIA is a defining moment on our journey to infinite inference, accessible to all," said Sid Sheth, founder and CEO at d-Matrix. "Being integrated into NVIDIA's latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform. That's the future d-Matrix has been building toward — ultra-low latency, energy-efficient inference XPUs and GPUs working together, at rack scale, to deliver premium AI experiences."

"NVLink Fusion enables partners to integrate custom silicon with NVIDIA's deep ecosystem of NVLink, advanced packaging, rack-scale systems and networking technologies," said Jensen Huang, founder and CEO of NVIDIA. "With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms — expanding accelerator choice for customers building the next generation of AI factories."

"Purpose-built connectivity is what turns innovative compute into high-performing AI factories," said Jitendra Mohan, CEO of Astera Labs. "Our partnership with d-Matrix and NVIDIA brings this vision to life within the NVLink Fusion ecosystem, delivering high-throughput for low latency AI inference."

The Rise of the Premium Token Economy 
As agentic AI workloads have spiked inference demand, AI service providers are increasingly seeking mixed-architecture systems to deliver inference economics their customers require.

The d-Matrix MGX rack system is designed for latency-sensitive applications such as AI coding assistants, real-time chatbots, and voice agents where interactivity is paramount and customers are willing to pay a premium for speed. Built on NVIDIA's mature, proven MGX ecosystem and supply chain, the system extends a unified rack architecture that gives AI factories the flexibility to deploy the right compute for each workload.

Using heterogeneous disaggregation, operators can split the workload between d-Matrix Raptor XPUs and NVIDIA Vera Rubin, allowing them to optimize each phase of inference. For one of today's most popular disaggregated applications, AI coding, GPUs can handle the compute-intensive prefill phase of a workload while d-Matrix inference XPUs speed up the latency-sensitive decode phase.

Raptor: d-Matrix's Next-Gen Inference XPU
A follow-on to the d-Matrix Corsair™ XPU platform currently in production, the d-Matrix Raptor platform extends d-Matrix's memory-centric architecture. Through a first-of-its-kind 3D DRAM stacking approach, Raptor brings a DRAM memory chip and an SRAM compute chip together to form a single "two-story" package.

Technical details about the 3D DRAM technology were recently published by IEEE and previewed by d-Matrix co-founder and CTO Sudeep Bhoja at the 2026 Hot Chips conference and .

d-Matrix has designed Raptor, which is expected to tape-out before the end of the year, from the ground up for integration with NVIDIA NVLink Fusion and the NVIDIA MGX rack-scale ecosystem, reflecting d-Matrix's commitment to building purpose-built inference silicon that works seamlessly alongside the NVIDIA stack.

Raptor is being actively evaluated at AI hyperscalers and frontier labs for its unique memory stacking solution and is backed by more than 100 patents.

Availability
Initial availability of d-Matrix Raptor XPUs integrated into the NVIDIA MGX rack is expected Q4 2027.