NVIDIA Vera Rubin NVL72 Leads MLPerf Inference v6.1 Benchmarks

NVIDIA claimed the top spot on MLPerf Inference v6.1 with its Vera Rubin NVL72, framing the results around tokens-per-watt efficiency — the metric that matters most for 'AI factories' where power, not silicon, is the binding constraint. Infrastructure engineers reacted to reported 2.5–3.7x throughput gains over the previous-generation GB300, a substantial generational leap if it holds in shipping hardware.
Alongside the benchmark, NVIDIA announced an alliance with Google and Emerald AI focused on energy-efficient data centers, expanded its CUDA-Q platform with CUDA-Q Logical for fault-tolerant quantum orchestration, and highlighted Perplexity's Portable Computer on Windows powered by RTX. The through-line is NVIDIA extending its moat from raw chips to a full stack spanning power efficiency, quantum orchestration, and local inference.
The most important caveat comes from the community itself: engineers flagged that Vera Rubin NVL72 is in 'preview' status, meaning buyers can't actually order it yet. That gap between benchmark leadership and availability is a recurring NVIDIA pattern — announce dominance early to freeze competitors' purchasing decisions while the product ramps.
Competitively, the timing is pointed. Samsung just co-led a $231M Series A for NVIDIA-rival inference-chip startup Euclyd, and Meta is deploying its in-house MTIA 450 chips — both signals that hyperscalers and investors are actively funding alternatives to reduce NVIDIA dependence. Vera Rubin's efficiency lead is NVIDIA's answer: make the incumbent hardware so power-efficient that custom silicon can't justify its switching cost. The skeptical read is that MLPerf results on preview hardware, tuned by the vendor, always flatter; real-world tokens-per-watt in customer data centers is the number that matters. Watch for general availability and independent efficiency measurements.