Nvidia's Vera Rubin platform promises 10x the token throughput per megawatt of its predecessor, giving the chipmaker a fresh edge as cloud giants race to cut AI energy costs.
Nvidia's Vera Rubin platform promises 10x the token throughput per megawatt of its predecessor, giving the chipmaker a fresh edge as cloud giants race to cut AI energy costs.

Nvidia Corp. unveiled expanded adoption of its Vera Rubin platform across four major cloud providers, with CoreWeave tests showing the system delivers 10 times the token throughput per megawatt compared with the prior-generation Grace Blackwell NVL72. Google Cloud, Microsoft Azure and Oracle Cloud have also begun deploying the Vera Rubin NVL72 system, Nvidia said July 21.
"We're on a road map to crank out new architectures, not just GPUs but CPUs," Ian Buck, Nvidia's vice president of accelerated computing, said during a technical workshop at the company's Santa Clara headquarters. "We're going to keep innovating, because it's do this or die in Silicon Valley."
The Vera Rubin NVL72 system integrates 36 Vera CPUs paired with 72 Rubin GPUs, along with NVLink 6 interconnects, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-6 switching. Nvidia said the platform uses a monolithic chip design — abandoning the chiplet architecture used by many modern processors — to reduce memory bandwidth penalties. The system is 100 percent liquid-cooled and designed as "cable-free compute," allowing installation time to shrink from hours to minutes, according to the company.
The efficiency gains come as Nvidia faces intensifying competition from Advanced Micro Devices Inc., which is expected to detail its Helios AI chip rack at its annual conference this week. Nvidia shares have benefited from the AI infrastructure buildout, and the Vera Rubin platform — shipping in the second half of 2026 with early customers including Microsoft Corp., OpenAI and Oracle Corp. — could help sustain that momentum.
Monolithic Design Breaks From Industry Norm
Nvidia's decision to build the Vera CPU as a single, monolithic die rather than stitching together multiple chiplets — an approach pioneered by AMD in its x86 processors — addresses a key bottleneck in AI workloads. Hannah Coutand, who runs Nvidia's CPU product marketing, argued that chiplet architectures impose "a heavy tax on memory bandwidth and data movement." The Vera CPU's localized memory subsystems offer nearly three times the memory bandwidth of the Grace Blackwell generation, a critical advantage as high-bandwidth memory remains in short supply.
The Vera CPU is built on the ARM architecture, an alternative to the x86 standard that dominates data center CPUs. Nvidia is selling the Vera CPU as a stand-alone product and has told Chinese customers these could be ready as soon as August, according to a WIRED report. The company said the Vera Rubin NVL72 system processes 10 times as many tokens per watt as Grace Blackwell, though the benchmarks used slightly older generations of competing AMD and Intel CPUs for comparison.
CoreWeave, the first cloud provider to complete system-level validation of Vera Rubin NVL72, has been a key early partner. Nvidia increased its investment in CoreWeave to $2 billion in January, and the company's backlog now approaches $100 billion, with more than 3.5 gigawatts of power under contract. The Vera Rubin platform already covers more than 350 factory nodes across 30 countries, Nvidia said.
The competitive stakes are high. AMD controlled a growing share of the data center CPU market entering 2026, and its Helios rack system is designed to challenge Nvidia's dominance in AI infrastructure. Nvidia's ability to deliver on Vera Rubin's production timeline is particularly sensitive after the previous-generation Blackwell chips faced overheating issues that forced design changes and delayed shipments.
For investors, the Vera Rubin ramp represents both opportunity and risk. Nvidia trades at roughly 35 times forward earnings, reflecting expectations that the company can maintain its dominant position as AI infrastructure spending grows. The global cloud AI market is projected to reach $780.6 billion by 2034, according to Fortune Business Insights, up from $133.4 billion in 2026. If Vera Rubin delivers on its efficiency claims, it could help Nvidia defend its roughly 80 percent share of the AI chip market against AMD's encroachment. If production delays or overheating issues resurface, the valuation premium could come under pressure.
This article is for informational purposes only and does not constitute investment advice.