NVIDIA announces Vera Rubin AI data center platform
NVIDIA announced its Vera Rubin platform on March 16, 2026, describing it as a new platform for agentic AI. The company said seven new chips are in full production and designed to scale AI factories.

Key takeaways · 4
- 01
NVIDIA says the seven new chips are in full production, while partner products based on Vera Rubin are expected from the second half of 2026.
- 02
The system combines 72 Rubin GPUs and 36 Vera CPUs connected by NVLink 6.
- 03
NVIDIA claims LPX paired with Vera Rubin can provide up to 35 times higher inference throughput per megawatt.
- 04
BlueField-4 STX uses dedicated KV-cache storage processing, which NVIDIA says can boost inference throughput by up to five times.
A platform built as one system
NVIDIA describes Vera Rubin as seven chips, five racks and one large supercomputer, with components intended to operate together across AI workloads.[1] The company says the platform supports pretraining, post-training, test-time scaling and agentic inference.[1] One system configuration integrates 72 Rubin GPUs and 36 Vera CPUs connected through NVLink 6.[1] NVIDIA also describes a system integrating 256 Vera CPUs for scalable, energy-efficient capacity.[1] The announcement frames the components as a unified platform rather than standalone chips.[1]
Performance claims for training and inference
NVIDIA says training large mixture-of-experts models with Vera Rubin takes one-fourth as many GPUs as with Blackwell.[1] For inference, the company claims up to 10 times higher throughput per watt at one-tenth the cost per token.[1] Those figures are NVIDIA’s comparisons, and they distinguish GPU requirements for training from throughput and cost claims for inference.[1] NVIDIA says the platform is optimized for trillion-parameter models and million-token context.[1]
Storage, networking and power
NVIDIA’s Groq 3 LPX rack has 256 LPU processors, 128GB of on-chip SRAM and 640 TB/s of scale-up bandwidth, according to the company.[1] NVIDIA says LPX combined with Vera Rubin can deliver up to 35 times higher inference throughput per megawatt.[1] BlueField-4 STX is described as AI-native storage that extends GPU memory across the POD; dedicated KV-cache processing can boost inference throughput by up to five times.[1] The system can use Spectrum-X Ethernet or Quantum-X800 InfiniBand switches.[1] NVIDIA says its deployment can fit 30% more AI infrastructure within a fixed-power data center.[1]
Availability and early customer signal
NVIDIA says Vera Rubin-based products from partners will begin to be available in the second half of 2026, and Groq 3 LPX racks are also scheduled for that period.[1] The company named Amazon Web Services, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure among leading cloud providers, and listed Cisco, Dell Technologies, HPE, Lenovo and Supermicro as manufacturers expected to deliver servers based on Vera Rubin products.[1] OpenAI CEO Sam Altman said OpenAI would use Vera Rubin to run more powerful models and agents at scale and deliver faster, more reliable systems.[1]
For infrastructure teams, the announcement offers a view of NVIDIA’s planned hardware and software stack, but the stated performance figures are company claims rather than deployment results in this evidence. Teams weighing adoption can use the second-half 2026 availability window and the system’s power and throughput claims to frame evaluation questions.
Why it matters
Test yourself on this story — 2 questions.
Create a free account to take the quiz, earn XP, and get a daily session built for your industry.
Take the quizHow this developed
5 October 2026
NVIDIA announces Vera Rubin AI data center platform
Sources
- NVIDIA Vera Rubin Opens Agentic AI Frontier | NVIDIA Newsroomnvidianews.nvidia.com
- NVIDIA unveils AI data center lineup at GTC 2026newsbytesapp.com