Agentic AI systems generate far more token activity than a single chat turn because each reasoning step—querying databases, invoking sub‑agents, synthesizing results—feeds its output back as the next input. This causes context to accumulate, often reaching hundreds of thousands of tokens, while simple chat or summarization tasks stay within the 1 K–8 K range. As a result, the infrastructure that runs these workloads must handle long, variable‑length sequences efficiently.
NVIDIA’s Vera Rubin NVL72 system was measured against the GB300 NVL72 platform using the SemiAnalysis AgentX workload, which captures real‑world agentic coding sessions with actual context growth, tool calls and sub‑agent spawning preserved. The benchmark shows up to 30x higher throughput per megawatt for Vera Rubin NVL72, meaning that for the same power envelope it can perform thirty times more agentic work. These results reflect the different performance profile of agentic workloads, where input lengths can far exceed those of conventional inference tasks.
- Vera Rubin NVL72 vs GB300 NVL72 comparison
- Throughput per megawatt: up to 30x higher (Vera Rubin)
- Workload: SemiAnalysis AgentX (real‑world agentic coding, preserved context growth, tool calls, sub‑agent spawning)
- Token demand: agentic workloads consume ~15x more tokens than a simple chat request
- Context size: can reach hundreds of thousands of input tokens, far beyond 1 K–8 K chat range
Why this matters
This efficiency gain suggests that power‑constrained AI factories can scale agentic pipelines without proportional increases in energy cost, directly addressing the token‑intensive nature of multi‑step reasoning. While the source notes the 30x throughput per MW figure, the implication is that operational expenditure for sustained agentic workloads could drop significantly, making continuous deployment more feasible.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
