Back to Newsroom

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

By Modelverse Editorial·August 24, 2026·2 min read
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

NVIDIA has moved the Vera Rubin rack‑scale system, which includes the newly production‑ready Groq 3 LPX accelerator, into full manufacturing. The system is positioned for agentic AI workloads that demand rapid token generation over long input contexts.

In an Artificial Analysis benchmark using the open‑source Gemma 4 31B model, the Vera Rubin + Groq 3 LPX combination produced 3,400 output tokens per second while handling a 100,000‑token context, a throughput four times higher than the closest competing platform. Early adopters such as SpaceXAI (Vera CPUs for next‑gen agentic AI), CoreWeave (Spectrum‑X Multiplane interconnect with parallel switches for lossless, high‑bandwidth networking), and Nebius (first cloud to offer Groq 3 LPX) illustrate the ecosystem forming around this architecture.

The Vera Rubin rack is built around a NVL72 configuration that pairs high‑density CPU cores with the Groq 3 LPX inference engine, while Spectrum‑X provides a flat Ethernet fabric that eliminates packet loss under heavy traffic. This end‑to‑end co‑design allows developers to run existing frameworks such as CUDA and TensorRT without modification, targeting workloads where agents exchange large context windows and produce sustained token streams.

  • 3,400 output tokens per second
  • 100,000‑token context window
  • 4× speed advantage over nearest alternative
  • Vera Rubin NVL72 + Groq 3 LPX + Spectrum‑X Multiplane

Why this matters

The result indicates that inference infrastructure is evolving to meet the specific demands of agentic systems, which generate many tokens and rely on expansive contexts. By tightly coupling compute, networking, and token‑generation hardware, NVIDIA’s extreme‑codesign approach reduces latency and improves throughput for long‑chain reasoning tasks. The benchmark figures and early partner deployments provide concrete evidence that the integrated stack can deliver measurable performance gains in real‑world agentic AI scenarios.

Share this article

Found this insightful? Share it with your community on Reddit, X, or copy the link.

ai-newsbriefnvidia

Footnotes & Primary References

Related content

Situational Awareness, star AI hedge fund that nearly imploded, now being probed by the SEC

The AI hedge fund went from "the talk of Wall Street" to "subject of federal subpoenas" faster than you can say "diversify your portfolio."

Read article

Wire It, Run It, Deploy It: AI Workflows in Gradio

Wire It, Run It, Deploy It: AI Workflows in Gradio

Read article

Scientific Data Analysis with LabPlot in Python: Signal Processing, Spectral Peak Fitting, Visualization, and Batch Automation

In this tutorial, we explore a LabPlot-inspired scientific data analysis workflow in Python while preserving the structure and terminology of LabPlot’s aspect tree, analysis kernel...

Read article