NVIDIA has announced that NVIDIA Groq 3 LPX, a platform designed for low-latency token generation, has entered full production, defining its role within the Vera Rubin architecture. The move primarily targets inference workloads tied to AI agents: systems that do not merely return an answer, but maintain very large contexts, call external tools, execute code, and coordinate multiple steps to complete a task.
The announcement, unveiled at Hot Chips in Palo Alto, does not introduce Vera Rubin as a standalone product. Instead, NVIDIA is proposing a package combining compute, inference acceleration, and networking, arguing that this integration will become decisive with the rise of reasoning-oriented models and multi-agent workflows. Under this blueprint, Vera Rubin NVL72 manages the rack-scale infrastructure, Groq 3 LPX targets the rapid response generation phase, and Spectrum-X Ethernet is designed to transport data between nodes in the AI factory.
The bottleneck shifts to token generation
In recent years, the industry’s focus has largely centered on training large models. However, for applications serving users and enterprises in real time, inference dictates response times, available capacity, and operational costs. The issue becomes more pronounced when a model must process documents, knowledge bases, or very long histories before it begins generating text, code, or instructions.
NVIDIA places Groq 3 LPX precisely at this final stage of the process. The stated goal is to reduce latency and increase the rate at which tokens are generated, while the Vera Rubin platform handles the heavier burden of context processing. This separation of duties is crucial for agents: a slow response can delay a chain of tool calls, data checks, code executions, and further reasoning steps.
As a performance indicator, NVIDIA cites an Artificial Analysis test using Gemma 4 31B, described as an open-source model for agentic use cases. In the benchmarked scenario, with a 100,000-token context, Groq 3 LPX reportedly reached 3,400 output tokens per second. According to the company, this figure is four times higher than the closest competing platform in the benchmark. While it provides a useful metric to frame the product’s positioning, it remains a result tied to a specific model, configuration, and workload; it does not automatically translate to the performance that will be seen across every application.
Vera Rubin as a coordinated architecture
NVIDIA’s message is that inference efficiency will not depend on a single accelerator. Serving models with very large context windows requires combining computing power with adequate memory and interconnects capable of rapidly moving vast amounts of data between racks and systems. Hence the emphasis on so-called codesign—the coordinated design of the entire infrastructure stack.
Vera Rubin NVL72 is the rack-scale cornerstone of this proposal. Groq 3 LPX is presented as a complement for requests where generation speed and responsiveness matter most, rather than as a replacement for the primary infrastructure. NVIDIA thus aims to offer cloud providers and enterprises a modular setup: resources for processing context, acceleration for output, and networking to tie everything together at data center scale.
The networking component is Spectrum-X Multiplane. CoreWeave has already begun deploying it in production to connect Vera Rubin racks across multiple parallel switches. NVIDIA claims the configuration enables a high-bandwidth, flat, and lossless network for AI workloads. In environments where a service distributes work across numerous nodes, the network is not a minor detail: delays and congestion in data transfers can diminish the benefits of ultra-fast accelerators.
Early adopters and the cloud front
Nebius will be the first AI cloud to adopt NVIDIA Groq 3 LPX. The company plans to integrate it into its Nebius Token Factory alongside Vera Rubin NVL72, providing developers with inference capacity tailored for interactive agents, coding tools, and services requiring real-time responses. The move to the cloud is significant because it makes this technology accessible even to those who do not plan to build rack-scale infrastructure in-house.
For NVIDIA, availability through a provider also serves as a real-world test. The announced metrics will need to translate into pricing per token, availability, quality of service, and compatibility with different models—factors that directly influence choices made by developers building agent-based products. The company touts lower token costs, but the published materials disclose no price lists, Nebius service rates, or detailed economic comparisons. For now, therefore, it is possible to note the platform’s trajectory, but not quantify the financial advantage for users.
Among the other cited partners is SpaceXAI, which plans to base its future AI architecture on Vera Rubin, from terrestrial data centers to orbital satellites. The company intends to use NVIDIA Vera CPUs for CPU-intensive tasks that accompany agent workloads, including orchestration, tool use, code execution, data processing, and simulation. For these CPUs, NVIDIA highlights per-core performance, memory bandwidth, and predictable behavior under load, aiming to keep GPU resources utilized efficiently.
A proposal that also raises the level of integration
Full production of Groq 3 LPX marks a transition from a technology showcase to industrial availability, but it does not eliminate the uncertainties typical of a new platform. Generation speed is only one of the variables defining an agent's experience: prompt processing time, the reliability of external software calls, model quality, memory management, and overall infrastructure costs also matter.
Furthermore, a tightly integrated architecture can offer advantages when all its components work together, but it requires companies to carefully evaluate interoperability, deployment models, and dependency on a single hardware and software supply chain. In its announcement materials, NVIDIA provided no details on geographic availability, complete commercial configurations, or expansion timelines beyond the specified partners.
The most significant takeaway from the move is therefore strategic. NVIDIA is adapting its lineup to a phase in which generative AI is used less as a simple conversational interface and more as an infrastructure capable of executing complex workflows. With Groq 3 LPX in production, Vera Rubin NVL72, Spectrum-X Multiplane, and Vera CPUs, the group is attempting to cover every segment of that pipeline. Production adoption by cloud providers and testing against real-world workloads—beyond published benchmarks—will reveal how effectively this blueprint translates into an operational advantage for AI agents.



