Designing a proprietary accelerator is no longer enough to truly enter the large-scale artificial intelligence market. For hyperscalers and AI-native companies, success depends on the ability to turn that chip into a seamless operational system: racks, power delivery, cooling, interconnects, software, monitoring, and maintenance must work as a single facility. NVIDIA aims to step in at this exact stage with NVLink Fusion, a platform designed to connect custom-developed XPUs to the company's infrastructure.
The offering comes at a stage where major cloud operators and specialized players are investing in their own AI semiconductors. XPUs—a broad term used to describe accelerators dedicated to specific workloads—can offer advantages when a company wants to optimize training or inference for its own models, software, and volumes. However, the economics of an AI factory are not measured by the chip in isolation: what matters are tokens produced per second, power consumption per token, service cost, hardware utilization rates, and the actual time the facility remains available.
In other words, specialized silicon risks losing part of its edge if placed in a platform plagued by bottlenecks in inter-accelerator communication, resource management, or system assembly. This is the premise behind NVLink Fusion: let customers design the XPU while providing mature components and integrations to complete the factory.
The value of the network between accelerators
At the core of the initiative is extending the NVLink domain to custom accelerators. In larger models, mixture-of-experts architectures, and agentic systems, workloads are distributed across many processors. Under these conditions, the speed at which chips exchange data can directly affect infrastructure utilization. If accelerators spend too much time waiting for information from other nodes, the investment in compute power is underutilized and the cost per token rises.
NVIDIA points to the sixth generation of NVLink as the foundation for the scale-up interconnect—that is, the high-speed network within a compute domain. The configuration cited by the company scales up to 72 XPUs and aims to combine high bandwidth and low latency. According to NVIDIA, end-to-end transfers between XPUs can achieve three times lower latency compared to standard Ethernet-based alternatives, while the packet rate can be ten times higher. These figures are provided by the vendor and should be viewed in the context of the benchmarked configurations, not as a universal metric applicable to every workload.
The roadmap outlined by NVIDIA also targets domains of up to 1,152 accelerators and the use of co-packaged optics—a solution where optical components are brought closer to or integrated into the silicon package to address the bandwidth and power consumption limits of traditional connections. The goal is to support increasingly dense systems without relying entirely on slower external links for scalability.
NVLink Fusion also includes NVLink-C2C, the interface capable of connecting an XPU to NVIDIA Vera CPUs or processors from other ecosystem players. NVIDIA claims up to six times greater energy efficiency for this solution compared to PCIe. Here again, the actual benefit will depend on the CPU type, platform design, and the intensity of data exchange between control and compute. A tight interconnect between these two components is particularly relevant for agentic applications, where task orchestration and model execution can require frequent communication.
From chip to rack, the most expensive stretch
The promise of a proprietary XPU is often associated with freedom from general-purpose hardware. In practice, however, the transition from chip design to data center deployment introduces a long list of industrial challenges. High-speed interfaces to CPU and network must be validated, switches and optical components sourced, compute and switching trays designed, rack-level power and thermal distribution verified, storage and security integrated, suppliers coordinated, and the system made serviceable without bringing down the entire installation.
This complexity also impacts timelines. A company building everything in-house retains greater architectural control, but must take on the risk of validating every link in the chain. NVIDIA is thus attempting to position NVLink Fusion as an industrial shortcut rather than a mere interconnect protocol: customers can adopt a tailored accelerator without rebuilding the entire rack-scale environment from scratch.
The package includes the MGX architecture, already used for NVIDIA-based systems, alongside its associated supply chain. MGX provides modular building blocks for racks, cooling, and power, and is also designated as the foundation for Vera Rubin NVL72 systems. The company also points to an evolution toward 800 VDC solutions, signaling how power density is becoming a central constraint in AI factory design.
Assembly standardization can carry significant weight. Jack Luoh, head of product and solution at QCT and Quanta Computer, explained that with Vera Rubin NVL72 the company is targeting an almost fully automated production line, and that part of the investments can be reused when an XPU leverages NVLink Fusion. For a system manufacturer, being able to adapt an already established supply chain means reducing variables during production ramp-up; for chipmakers, it means avoiding the need to build equivalent-scale manufacturing and integration capabilities on their own.
An opening that keeps NVIDIA at the center
NVLink Fusion represents an opening toward non-NVIDIA silicon, but it is not equivalent to a neutral platform. The value proposition stems precisely from gaining access to the NVIDIA ecosystem: the NVLink interconnect, MGX architecture, software, and associated suppliers. Those adopting this model can choose a CPU and develop an XPU tailored to their needs, but they place that choice within a significant portion of the Californian company's stack.
Tim Wilson, vice president and general manager of data center silicon engineering at Intel, emphasized that the initiative would allow customers to choose CPU architecture, performance tiers, and software capabilities based on their workloads. The reference to Intel highlights NVIDIA's push to expand the perimeter of partners around its infrastructure, rather than limiting itself to selling systems built exclusively from its own GPUs and CPUs.
For hyperscalers, the choice will come down to a trade-off. Building a fully in-house environment can ensure differentiation and control, but it demands capital, expertise, and very long lead times. Relying on an established platform can shorten the path to deployment and reduce operational risk, at the cost of greater dependence on NVIDIA standards, components, and roadmaps. AI-native companies, which often lack the bargaining power or organizational scale of major cloud providers, may find the second route more compelling.
It remains to be seen how quickly the ecosystem will translate into commercial products, which XPUs will actually be connected to NVLink Fusion, and their real-world performance across various inference and training scenarios. NVIDIA has outlined the architecture and expected benefits, but the ultimate test will be partners' ability to bring complete, reliable, and cost-per-token competitive systems to market. In an AI factory, the chip is merely the starting point: the final outcome depends on how effectively the entire chain delivers continuous compute.



