Key Takeaways: NVIDIA's first on-chip benchmarks for Vera Rubin NVL72 show a 30x throughput-per-megawatt gain over GB300 NVL72 on real agentic coding workloads.
Key Takeaways: NVIDIA's first on-chip benchmarks for Vera Rubin NVL72 show a 30x throughput-per-megawatt gain over GB300 NVL72 on real agentic coding workloads.

NVIDIA's Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72 on real agentic coding workloads, per first on-chip benchmarks disclosed at Hot Chips 2026.
"Inference is the growth engine of AI," Jensen Huang, founder and CEO of NVIDIA, said. "Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI."
The benchmarks, run on SemiAnalysis's AgentX workload using DeepSeek-V4-Pro (1.6T parameters), replay production-style coding agent sessions that accumulate context, invoke tools, and spawn sub-agents — a pattern consuming up to 15x more tokens than ordinary chat, per OpenRouter data. GB300 NVL72 itself delivers up to 15x higher throughput per megawatt than H200 NVL8 on the same workload, and up to 80x on larger mixture-of-experts models like Kimi K3 2.8T. The gains come from system-level optimizations: MoE serving runtimes (SGLang, TensorRT-LLM, vLLM), DeepGEMM-based kernels, mixed-precision formats (MXFP4, MXFP8), and NVIDIA Dynamo's session-aware serving stack with KV-cache-aware routing.
The results reframe AI infrastructure economics. With power and data center capacity the binding constraint on AI expansion, throughput per megawatt — not raw peak compute — is becoming the metric that determines AI factory profitability. NVIDIA also announced full production of Groq 3 LPX inference accelerators, which hit 3,400 tokens per second on Gemma 4 31B at 100K context, and revealed SpaceXAI is deploying Vera CPUs for gigawatt-scale compute, including plans for orbital AI.
The Groq 3 LPX, built from technology NVIDIA acquired from Groq for roughly $20 billion in December 2025, is a dedicated interactive inference accelerator designed to complement Vera Rubin NVL72. Each 2U liquid-cooled tray carries sixteen Groq 3 LPUs with 500MB of on-chip SRAM, a host CPU, and either a BlueField-4 DPU or ConnectX-9 NIC. A rack-scale deployment integrates up to 256 LP30 accelerators.
In Artificial Analysis benchmarks running Gemma 4 31B with a 100,000-token context window, Groq 3 LPX reached 3,400 output tokens per second — the fastest result ever recorded for that model and roughly 4x faster than the nearest alternative platform. For latency-sensitive agentic coding, generating 5,000 tokens drops from about 50 seconds to 1.5 seconds. NVIDIA's architecture splits the work: Rubin GPUs handle large-scale context processing and prefill, while LPX accelerates latency-sensitive decode.
The comparison warrants scrutiny. Cerebras's CS-3 accelerator achieves 882 tokens per second on the same Gemma 4 31B test using just one to two chips, while NVIDIA requires at least 64 Groq 3 LPUs. Cerebras also unveiled its CS-4 accelerator last week, based on the WSE-3T wafer-scale engine with 2x the compute and memory bandwidth of its predecessor.
Nebius is the first AI cloud partner to deploy Groq 3 LPX, integrating it into its Nebius Token Factory production inference service. "Generation is the phase of inference that determines how responsive an AI system actually is," Danila Shtan, chief technology officer of Nebius, said. Groq, the inference provider, is expected to be among the earliest adopters.
NVIDIA also disclosed that its Vera CPU — 88 custom Olympus cores with LPDDR5X memory delivering 1.2TB/s of bandwidth — is entering full production. The chip is designed for the CPU-intensive work agents generate between GPU inference calls: tool execution, Python code, context retrieval, and data processing. NVIDIA claims Vera completes agentic AI, reinforcement learning, and data processing tasks up to 1.8x faster than x86 CPUs.
SpaceXAI has begun scaling deployment of Vera CPUs and is building a gigawatt-scale compute factory on the Vera Rubin platform. The company plans to deploy an optimized version of Vera Rubin NVL72 in its first-generation Starmind AI satellites, with mass deployment targeted for 2028.
NVIDIA's announcements show the competitive battleground in AI infrastructure shifting from training performance to token generation efficiency, energy consumption, and cost per useful output. AMD and Cerebras are making aggressive moves in low-latency inference, and benchmark comparisons between NVIDIA and Cerebras have drawn scrutiny over chip-count differences. NVIDIA reports earnings this Wednesday, where AI inference demand and Groq commercialization progress are expected to be focal points alongside Blackwell and Vera Rubin order momentum.
This article is for informational purposes only and does not constitute investment advice.