
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference V6.1 Debut
AI Executive Summary
NVIDIA released preview MLPerf Inference v6.1 benchmark results for its Vera Rubin NVL72 platform, showcasing up to 3.7x higher throughput on Qwen3-VL using vLLM and NVIDIA Dynamo, and 2.5x higher on DeepSeek-R1 using TensorRT-LLM compared to the GB300 NVL72.
The system leverages sixth-generation NVLink and NVLink Switch hardware, NVFP4 precision, and disaggregated serving to achieve 30x better performance on agentic workloads like SemiAnalysis AgentX.
Nebius also submitted competitive preview results utilizing the same hardware architecture.
Why It Matters
Strategic TakeawayFull-stack co-design combining NVFP4 precision, sixth-generation NVLink interconnects, and specialized software libraries drastically improves token generation economics for complex reasoning and multimodal models. This hardware-software integration redefines rack-scale performance boundaries for multi-step AI agent workflows.
Multi-Vector Implications
- TECHNICALImplementation of sixth-generation NVLink and NVLink Switch delivers 10x higher packet rates and 3x lower latency, enabling efficient disaggregated prefill/decode serving at rack scale.
- MARKETDeployment of Vera Rubin NVL72 shifts data center unit economics by multiplying token generation capacity per rack, directly eroding the cost-per-token baseline established by GB300 systems.
- GOVERNANCEAdoption of NVFP4 low-precision formats requires strict validation protocols to ensure multi-modal and reasoning model output quality remains uncompromised across production workloads.
Strategic Outlook
12-18M HorizonOver the next 12-18 months, cloud service providers and tier-1 infrastructure vendors like Nebius will aggressively transition production clusters to Vera Rubin NVL72 architectures to capture multi-fold throughput gains in reasoning and agentic workloads. Software frameworks will increasingly mandate tight coupling with TensorRT-LLM and NVIDIA Dynamo to exploit native NVFP4 and disaggregated serving feature.
Referenced Coverage & Sources
Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS - competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa.
Physical AI Takes the Wheel: How the World's Robotaxi Leaders Are Building with NVIDIA Technologies
The global robotaxi market - physical AI's first commercial breakthrough - is projected to reach $400 billion by 2035, with over 6 million commercial.
Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers
AI factories are the infrastructure of the intelligence era.
Inference
Inference is the process of using a trained AI model to make predictions or generate text based on new inputs. During inference, data flows forward through the neural network to produce an output, without modifying the model's weights.
NVIDIA
NVIDIA is a pioneer of GPU computing, dominating the hardware market for AI acceleration, training, and inference with its high-performance Hopper and Blackwell architectures.
Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:
Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.