The Information Machine

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

NVIDIA Blog · NVIDIA Writers · 2026-07-21

NVIDIA announces the Vera Rubin NVL72 platform entering full production across CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud, with CoreWeave benchmarks showing 10x more tokens per megawatt than the prior Grace Blackwell generation on DeepSeek-R1.

Open original ↗

Appears in

Extraction

Topics: ai-hardwaregpu-infrastructureinference-efficiencysovereign-ai

Claims

  • CoreWeave's DeepSeek-R1 benchmark on Vera Rubin NVL72 shows 10x improvement in tokens per second per megawatt compared with Grace Blackwell NVL72.
  • The NVIDIA Vera CPU delivers 2x single-threaded performance and supports up to 1.6x more concurrent AI agents than competing CPU designs.
  • Vera Rubin NVL72's 45-degree Celsius liquid cooling inlet enables chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually.
  • Microsoft and Mistral signed a multibillion-dollar agreement to expand European AI infrastructure using tens of thousands of Vera Rubin GPUs, targeting sovereign AI requirements.
  • Google Cloud A5X instances powered by Vera Rubin NVL72 deliver up to 10x lower inference cost per token and 10x higher token throughput per megawatt versus prior generation.

Key quotes

CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72 — landing directly on the metric that matters most for power-constrained AI factories.
Tokens per megawatt is the metric that determines whether AI infrastructure can profitably scale.
Sovereign AI should not force organizations to choose among innovation, economics and control.