NVIDIA has moved the Vera Rubin rack‑scale system, which includes the newly production‑ready Groq 3 LPX accelerator, into full manufacturing. The system is positioned for agentic AI workloads that demand rapid token generation over long input contexts.
In an Artificial Analysis benchmark using the open‑source Gemma 4 31B model, the Vera Rubin + Groq 3 LPX combination produced 3,400 output tokens per second while handling a 100,000‑token context, a throughput four times higher than the closest competing platform. Early adopters such as SpaceXAI (Vera CPUs for next‑gen agentic AI), CoreWeave (Spectrum‑X Multiplane interconnect with parallel switches for lossless, high‑bandwidth networking), and Nebius (first cloud to offer Groq 3 LPX) illustrate the ecosystem forming around this architecture.
The Vera Rubin rack is built around a NVL72 configuration that pairs high‑density CPU cores with the Groq 3 LPX inference engine, while Spectrum‑X provides a flat Ethernet fabric that eliminates packet loss under heavy traffic. This end‑to‑end co‑design allows developers to run existing frameworks such as CUDA and TensorRT without modification, targeting workloads where agents exchange large context windows and produce sustained token streams.
- 3,400 output tokens per second
- 100,000‑token context window
- 4× speed advantage over nearest alternative
- Vera Rubin NVL72 + Groq 3 LPX + Spectrum‑X Multiplane
Why this matters
The result indicates that inference infrastructure is evolving to meet the specific demands of agentic systems, which generate many tokens and rely on expansive contexts. By tightly coupling compute, networking, and token‑generation hardware, NVIDIA’s extreme‑codesign approach reduces latency and improves throughput for long‑chain reasoning tasks. The benchmark figures and early partner deployments provide concrete evidence that the integrated stack can deliver measurable performance gains in real‑world agentic AI scenarios.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
