Technology

NVIDIA Vera Rubin NVL72 Hits 10x Tokens Per Megawatt in First Benchmarks

Early benchmarks show NVIDIA's Vera Rubin NVL72 delivering 10x more tokens per megawatt than Grace Blackwell, with CoreWeave, Google Cloud, and DeepInfra validating performance.

NVIDIA's Vera Rubin NVL72 platform achieved 10x better tokens-per-megawatt efficiency than Grace Blackwell in CoreWeave's DeepSeek-R1 benchmark. The rack-scale system uses extreme codesign across seven chips, with live deployments already underway at CoreWeave, Google Cloud, and DeepInfra, and serves as the foundation for Microsoft and Mistral's European sovereign AI infrastructure.

NVIDIA's next-generation Vera Rubin platform is beginning to show measured performance numbers from live hardware, and the early results are significant. CoreWeave, the first AI cloud to validate Vera Rubin NVL72, ran a DeepSeek-R1 benchmark and recorded a 10x improvement in tokens per second per megawatt compared to Grace Blackwell NVL72.

Tokens per megawatt has become the critical metric for AI infrastructure economics. More tokens per megawatt means either more intelligence from the same power budget, or the same workload on substantially less energy. For power-constrained data centers, this directly determines whether AI scaling remains profitable.

Architecture and Design

The gains come from what NVIDIA calls "extreme codesign" across seven chips and five rack trays: Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX. The system is engineered as a single unified architecture rather than assembled from separate off-the-shelf products.

At the center is the Vera CPU, built on a custom Olympus core that NVIDIA says delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency versus competing chiplet designs. NVLink 6 provides 260 TB/s all-to-all GPU communication, allowing the rack to behave as a single unified accelerator. This is especially important for mixture-of-experts models like DeepSeek-R1, where each token must be routed across distributed expert sub-networks.

The physical design also reflects lessons from three generations of rack-scale codesign. The Vera Rubin NVL72 system contains no cables, fans, or hoses in the compute tray, cutting assembly time from hours to one minute. A 45-degree Celsius liquid cooling inlet temperature enables chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually.

Partner Deployments

CoreWeave is deploying the NVIDIA Spectrum-X Ethernet SN6600-LD as its switching fabric, built on the 102.4 Tb/s Spectrum-6 switch chip with liquid cooling. The company provides a fully non-blocking, multi-plane, multi-rail spine and leaf fabric connecting Vera Rubin NVL72 GPUs without oversubscription.

Google Cloud has launched its first A5X instance on Vera Rubin NVL72, already in use by London startup Ineffable Intelligence for reinforcement learning-based "superlearner" systems. The A5X instances use NVIDIA ConnectX-9 SuperNICs combined with Google Virgo networking, enabling clusters that can scale to tens of thousands of GPUs within a single site and nearly a million across multi-site configurations.

DeepInfra, which processes nearly 5 trillion tokens per week with about 30% driven by agentic systems, independently benchmarked the Vera CPU and found it supports 1.6x more concurrent AI agents at the same quality of service and delivers 2.2x faster orchestration than alternative CPUs.

Sovereign AI in Europe

Vera Rubin is also the computing foundation for an expanded Microsoft-Mistral partnership focused on European AI infrastructure. Mistral is adding GPU capacity using thousands of Vera Rubin GPUs, with Mistral Medium 3.5 and OCR 4 now available in Microsoft Foundry. The partnership targets sovereign-ready AI across public cloud, cloud-connected, and fully disconnected private cloud environments.

Compared with NVIDIA GB200 NVL72, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens. For organizations weighing AI infrastructure investments, these are not marginal improvements. They are the difference between scaling affordably and hitting a power wall.