AI inference chipmaker d-Matrix is adopting NVIDIA’s NVLink Fusion technology to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform, marking a move towards large-scale deployment of specialised inference processors.
The integration will allow d-Matrix to connect Raptor XPUs through NVIDIA NVLink scale-up technology and Spectrum-X scale-out networking, while leveraging NVIDIA’s MGX rack architecture and broader AI platform.
According to d-Matrix, the approach is designed to provide a faster and lower-risk path from custom silicon development to rack-scale deployment, allowing its customers to deploy and scale AI inference infrastructure without having to develop an entirely separate infrastructure stack.
“Demand for inference is soaring, but capital, time and energy remain finite,” said Sid Sheth, co-founder and CEO of d-Matrix. “With NVLink Fusion and MGX, we can integrate our Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference.”
NVIDIA’s NVLink Fusion is designed to allow third-party silicon companies to integrate custom XPUs and CPUs into NVIDIA’s infrastructure ecosystem. This enables chip developers to focus on processor architecture while taking advantage of established networking, rack infrastructure, power, cooling and supply-chain capabilities.
For d-Matrix, the move is significant because developing an XPU is only one part of building an AI infrastructure platform. Large-scale deployment also requires high-speed connectivity, networking, rack design, power management, cooling and system validation.
By adopting NVLink Fusion, d-Matrix can leverage the NVIDIA MGX ecosystem and its validated rack designs and infrastructure. The companies say a common rack architecture could allow data centres to support different processor types, including GPUs, CPUs and XPUs, without requiring completely separate rack designs.
Under its planned architecture, d-Matrix will use NVIDIA NVLink to connect its XPUs within a high-bandwidth, low-latency scale-up domain. The company also plans to integrate NVIDIA Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-X Ethernet networking.
The Raptor-based systems are also expected to operate alongside NVIDIA GPU-based platforms, including the Vera Rubin NVL72, supporting disaggregated AI inference workloads.
The development highlights a broader shift in the AI infrastructure market, where specialised accelerators are increasingly being deployed alongside general-purpose GPUs rather than operating as completely separate systems.
For NVIDIA, NVLink Fusion expands its AI infrastructure ecosystem beyond its own processors by allowing specialised silicon providers to tap into its networking, systems, software and rack-scale infrastructure.
For d-Matrix, the partnership could help shorten the path from chip development to commercial deployment while giving customers greater flexibility in selecting compute architectures based on workload requirements, performance, power efficiency and cost.
As demand for AI inference continues to grow, the ability to deploy different types of processors within a common infrastructure architecture could become increasingly important for data centres seeking to optimise performance and cost per token.








Leave a Reply