Skip to main content

NVIDIA Vera Rubin NVL72 Debuts with Major Performance Gains in MLPerf Inference v6.1

17 SEPTEMBER 2026·2 MIN READ·2 SOURCES·Official source plus independent coverage

NVIDIA's Vera Rubin NVL72 system has debuted in MLPerf Inference v6.1 preview submissions, delivering up to 3.7x better throughput than the GB300 NVL72.

NVIDIA Vera Rubin NVL72 Debuts with Major Performance Gains in MLPerf Inference v6.1

Key takeaways · 3

  • 01

    Vera Rubin NVL72 offers up to 3.7x better throughput than GB300 NVL72.

  • 02

    A 288-GPU submission using four GB300 NVL72 racks achieved 99% scaling efficiency.

  • 03

    Software optimizations in MLPerf Inference v6.1 provided up to 1.6x higher performance over v6.0.

New Benchmark Milestones

In its first MLPerf Inference preview submission, the NVIDIA Vera Rubin NVL72 system delivered up to 3.7x better throughput than the GB300 NVL72. [2] NVIDIA submitted these preview results on two demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL. [2] Using vLLM with the NVIDIA Dynamo open source inference framework, the Vera Rubin NVL72 delivered up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios. [2] On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput was up to 2.5x higher than the GB300 NVL72. [2]

Scaling and Software Gains

A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, demonstrating nearly linear throughput growth from a single-rack baseline. [2] Additionally, software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions resulted in up to 1.6x higher performance compared to v6.0. [2] Meanwhile, AI agents are expanding from cloud data centers to vehicles, robots, and other edge devices. [1]

What it means

The substantial throughput gains of the Vera Rubin NVL72, particularly its 3.7x advantage over the GB300 NVL72 on the Qwen3-VL benchmark, underscore rapid hardware advancements in handling demanding models. The near-linear scaling efficiency of the GB300 NVL72 suggests strong multi-rack operational viability. What the sources don't address: specific pricing or availability dates for the Vera Rubin NVL72 system.

Significant increases in inference throughput and scaling efficiency directly impact the economics of deploying large-scale AI models.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 17 September 2026

    NVIDIA Vera Rubin NVL72 Debuts with Major Performance Gains in MLPerf Inference v6.1

  2. 17 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.