NVIDIA Vera Rubin NVL72 Promises 30x Efficiency Gain for Agentic Workloads
NVIDIA reports its new Vera Rubin NVL72 systems deliver up to 30 times higher throughput per megawatt than the GB300 NVL72 on agentic AI workloads.

Key takeaways · 3
- 01
Agentic AI workloads consume 15x more tokens than a simple chat request.
- 02
Vera Rubin NVL72 delivers 30x higher throughput per megawatt than GB300 NVL72.
- 03
The new systems offer a 35x lower token cost for agentic tasks.
The Token Cost of Agents
According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. [1]
Agents and sub-agents keep reasoning until the task is done, driving increased token demand and making long-context handling central to agentic AI performance. [1] In agentic sessions, context accumulates across steps and can reach hundreds of thousands of input tokens. [1]
Vera Rubin Performance
New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. [1]
NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions. [1] For power-constrained AI factories, that translates directly into 30x more agentic work for the same energy footprint. [1] The Vera Rubin NVL72 also boasts a 35x lower token cost. [1]
What it means
The shift from simple chat interfaces to multi-step agentic workflows is creating immense pressure on computing infrastructure due to exponential token growth. By delivering up to 30x better power efficiency than its own GB300 NVL72, NVIDIA's Vera Rubin architecture directly targets the energy bottlenecks of deploying agents in power-constrained data centers. The reliance on the SemiAnalysis AgentX benchmark highlights an industry transition toward evaluating hardware on real-world tool calls and context accumulation rather than static summarization tasks. What the sources don't address: exactly when these Vera Rubin systems will be generally available to enterprise customers for deployment.
As AI usage shifts from static chats to autonomous agents, computing infrastructure must handle massive context accumulation. Extreme improvements in performance per watt will dictate which organizations can afford to deploy agentic architectures at scale.
Why it matters
Turn this story into practical AI skill after launch.
Get the release link for daily sessions built around your role and industry.
Join the waitlistHow this developed
25 August 2026
NVIDIA Vera Rubin NVL72 Promises 30x Efficiency Gain for Agentic Workloads
25 August 2026
Event created from source cluster.