d-Matrix has selected NVIDIA NVLink Fusion as the integration foundation for Raptor, the next generation of its XPUs designed for AI inference. The agreement goes beyond the simple adoption of a high-speed interconnect: the goal is to integrate d-Matrix’s specialized processors into NVIDIA’s infrastructure platform, leveraging the same rack architecture, networking, and part of the ecosystem that currently support the group’s GPU systems.
For a company designing silicon dedicated to model execution, the move is significant primarily on an industrial level. An accelerator may be competitive in terms of latency, power consumption, or execution cost, but making it into a data center requires much more: chip-to-chip interconnects, servers, power delivery, cooling, software, rack validation, and a supply chain capable of producing and delivering complete systems. d-Matrix will aim to avoid rebuilding this from scratch by plugging into NVIDIA’s infrastructure.
Raptor enters the NVLink domain
NVLink Fusion is NVIDIA’s initiative to open its scale-up interconnect to third-party CPUs and XPUs. Scale-up refers to tight, high-bandwidth communication between accelerators within a system or rack—a layer distinct from the network linking multiple racks across the data center. Under the announced plan, d-Matrix’s Raptor XPUs will be connected via NVLink in a single low-latency, high-bandwidth domain.
Completing the picture will be Spectrum-X, NVIDIA’s Ethernet platform designed for scale-out—connecting distinct systems and racks. d-Matrix has also indicated plans to integrate NVIDIA Vera CPUs, ConnectX-9 SuperNICs, and BlueField-4 DPUs. These components serve different roles: general-purpose CPUs, network interfaces, and processors for infrastructure and security services. The envisaged combination thus places Raptor within a heterogeneous environment, rather than an isolated appliance dedicated exclusively to the company’s chips.
Compatibility with the NVIDIA MGX rack architecture is the second piece of the announcement. MGX provides modular, pre-validated designs for building servers and racks tailored to AI workloads. In this configuration, an operator could adopt common chassis and infrastructure for GPUs, CPUs, and XPUs, rather than creating a dedicated layout for each processor family. It is a less visible aspect than interconnectivity, but often decisive when a project must scale from prototype to deployment across many systems.
The bet is on inference, not training
d-Matrix operates in the inference segment, the phase where an already trained model is used to generate responses, classify data, or perform application tasks. Demand for capacity in this phase has surged with the expansion of generative models and their use in services aimed at consumers and enterprises. It is also a market where metrics such as latency, energy required per computation, and cost per token carry direct weight.
The company presents Raptor as a processor specialized for these constraints. NVIDIA, for its part, offers d-Matrix a path to bring it into larger-scale infrastructures and deploy it alongside GPU-based systems, including those in the Vera Rubin NVL72 family mentioned in the plan. The premise is a disaggregated deployment of inference: different resources can be assigned to different tasks while maintaining a common foundation for racks and networking.
This does not imply that Raptor will replace NVIDIA GPUs. Rather, the point of the integration is to make a third-party accelerator usable within the technical and commercial perimeter of the NVIDIA ecosystem. For companies operating AI factories—the term NVIDIA uses to describe infrastructures dedicated to generating model outputs—the appeal lies in the ability to choose a specialized component without completely overhauling the rest of the platform.
NVIDIA claims that the sixth generation of NVLink can deliver up to 3 TB/s of all-to-all bandwidth per XPU, along with lower inter-XPU latency compared to standard Ethernet and higher packet rates. These are company-provided figures and should be viewed in the context of the actual configuration and workloads: the performance a customer will achieve with Raptor will depend on the final silicon, software, rack topology, and the type of models being run.
For NVIDIA, openness remains under platform control
The announcement clearly illustrates the evolution of NVIDIA's strategy. On the one hand, the company maintains a tightly integrated offering encompassing processors, networking, software, and rack design. On the other hand, with NVLink Fusion, it seeks to turn that very offering into a platform for third-party CPU and accelerator manufacturers. Rather than letting specialized chips build their own full stack, NVIDIA offers to host them within its own.
The distinction is also significant for d-Matrix's positioning. Tapping into an already widespread architecture can shorten time-to-market and reduce integration risks, but it also entails a dependence on NVIDIA's specifications, product cycles, and ecosystem. NVLink Fusion is therefore a technical opening that simultaneously reinforces NVIDIA's central role in defining the AI rack.
The program already brings together names from multiple tiers of the supply chain. NVIDIA cites, among others, AWS, Arm, Intel, Fujitsu, SiFive, Alchip, Astera Labs, GUC, Marvell, MediaTek, Samsung, Cadence, Synopsys, Ayar Labs, and Lightmatter. The presence of chip designers, IP providers, packaging firms, and interconnect specialists indicates that competition is not taking place solely around the individual accelerator. The ability to integrate it physically and logically into systems that can be built, cooled, and managed at scale matters more and more.
What the project is still missing
d-Matrix has announced the adoption of NVLink Fusion and the planned integration for Raptor, but did not disclose in the released materials a commercial availability date for the systems, pricing, finalized rack configurations, or independent benchmarks. As a result, the concrete advantages over solutions based on GPUs or other inference accelerators cannot yet be measured.
Software aspects also remain to be verified. In a mixed deployment, the quality of integration depends on the ability to orchestrate models, memory, networking, and workloads across different components. NVIDIA includes access to its software platform in its NVLink Fusion messaging, but d-Matrix will have to demonstrate how easy it will be for customers and developers to leverage Raptor in real-world workflows.
The next step will therefore be moving the agreement from the architectural level to commercially available products. If the integration hits the market on the anticipated terms, d-Matrix will be able to offer a specialized alternative for inference without asking data centers to adopt a parallel platform. For NVIDIA, meanwhile, the d-Matrix case will serve as a test of its ability to turn NVLink Fusion into a de facto standard for non-NVIDIA chips hosted in NVIDIA racks.



