Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Nvidia Vera Rubin NVL72 Targets 30x Tokens Per Watt

Nvidia Vera Rubin NVL72 Targets 30x Tokens Per Watt

Nvidia detailed the Vera Rubin NVL72 rack at Hot Chips 2026: 72 GPUs, 2 ZFLOPS of NVFP4 inference and up to 30x more tokens per megawatt.

Dr. Nova Chen
Dr. Nova ChenAug 26, 20266 min read

When Nvidia first unveiled the Vera Rubin architecture at GTC in March, the pitch was density: roughly three to four times the AI compute of Blackwell in a comparable power envelope. At Hot Chips 2026 on August 24, the company filled in what that looks like at rack scale, and the metric it chose to lead with is telling. Not FLOPS. Tokens per megawatt.

  • The Vera Rubin NVL72 rack holds 72 GPUs and is rated at 2 ZFLOPS for NVFP4 inference and 1.4 ZFLOPS for NVFP4 training
  • A full AI factory configuration carries 11 petabytes of HBM4 with 800 petabytes per second of memory bandwidth inside a 100MW power budget
  • Each GPU gets 3.6 TB/s of all-to-all NVLink 6 bandwidth, with 800 VDC power delivery and 45C liquid cooling
  • Nvidia claims up to 30x more tokens per megawatt than prior generations, plus a 13% peak power reduction during LLM training

Why Tokens Per Megawatt Became the Headline Number

Data centers are increasingly constrained by the power contract, not the purchase order. Once a site is capped at a given megawatt figure, the only way to serve more users is to extract more useful output from each watt — which makes tokens per megawatt the metric that actually maps to revenue.

Agentic workloads sharpen that pressure. Nvidia cites OpenRouter data showing agentic requests consuming up to 15x more tokens than a standard chat turn, because an agent that plans, calls tools and re-reads its own output generates far more traffic per user-visible result. ServeTheHome's coverage of the talk shows Nvidia plotting rack generations along a tokens-per-megawatt axis against interactivity, with Vera Rubin landing well above Blackwell NVL72 and far above the Hopper NVL8 generation. On an AgentX workload using a 140K-plus context, the Vera Rubin curve pulls above GB300 NVL72 near the 60 million total-token mark.

What Nvidia Announced Alongside the Rack

The same day, Nvidia detailed the pieces that sit around the NVL72. The Groq 3 LPX inference accelerator, co-designed with Vera Rubin, entered full production, with Nvidia citing an Artificial Analysis measurement of 3,400 output tokens per second on 100,000-token long-context work using Gemma 4 31B. Nebius is named as the first AI cloud adopting it.

Vera CPUs handle the orchestration side — tool use, code execution, data processing and simulation — the unglamorous work that dominates an agent's wall-clock time. And Spectrum-X Multiplane extends the hardware-accelerated Ethernet architecture to as many as 512,000 GPUs, with Nvidia claiming 1.6x better AI networking performance than off-the-shelf Ethernet and roughly 90% bandwidth retention through a single-plane failure. CoreWeave has it in production.

Is the Serviceability Story the Real Upgrade?

Possibly, and it is the part that gets least attention. The NVL72 design emphasises cable-free trays, hot-swappable components, adaptive sparsity optimisation and zero-downtime health checks. At 100MW, a rack that has to be taken offline to service is not a maintenance inconvenience — it is a measurable revenue event.

This is the same theme running through recent infrastructure work: the interesting gains increasingly come from keeping expensive silicon busy rather than from the silicon itself. Nvidia's own KV cache transfer work that skips seven-second re-prefills makes the point at the software layer, and Cerebras packing three wafer-scale chips into a rack makes it at the packaging layer.

For readers tracking where inference costs go next, the Hot Chips disclosures are worth more than the launch keynote was — they show the actual power, cooling and network assumptions behind the marketing curve. More on model and infrastructure news is on our AI page.

Sources: ServeTheHome — August 24, 2026; NVIDIA Blog — August 24, 2026; VideoCardz — August 2026.

More AI Stories