
CoreWeave Leads Cloud Providers in MLPerf® Inference V6.1 Performance with NVIDIA Blackwell Ultra
AI Executive Summary
CoreWeave achieved leading cloud provider inference throughput in MLPerf Inference v6.1 using NVIDIA Blackwell and Blackwell Ultra platforms.
Operating within the Datacenter Closed division's Available category across four NVIDIA platforms and four model families, CoreWeave utilized an NVIDIA GB300 NVL72 rack running the Qwen3-VL-235B-A22B model to sustain 1,196 queries per second in the server scenario and 1,135 samples per second in the offline scenario.
Why It Matters
Strategic TakeawayScaling multimodal and mixture-of-experts reasoning model requires tightly co-designed rack-scale hardware and serving software to compress inference costs and overcome memory bandwidth bottlenecks.
Multi-Vector Implications
- TECHNICALContinuous optimization of serving stacks, scheduling algorithm, and operations unlocks higher throughput from existing NVIDIA GB300 NVL72 hardware without waiting for next-gen silicon.
- MARKETHigh per-GPU throughput metrics in standardized benchmarks like MLPerf v6.1 give cloud providers a direct competitive advantage in lowering customer infrastructure costs for frontier-scale models.
- GOVERNANCEStandardized MLPerf evaluations provide transparent, auditable performance baselines that enterprise buyers must mandate when validating cloud AI infrastructure SLAs for production workloads.
Strategic Outlook
12-18M HorizonOver the next 12 to 18 months, cloud AI infrastructure competition will increasingly hinge on rack-scale software co-design and Day-1 availability of Blackwell Ultra systems to efficiently serve multimodal and MoE reasoning model at scale.
Referenced Coverage & Sources
Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference V6.1 Debut
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics.
Physical AI Takes the Wheel: How the World's Robotaxi Leaders Are Building with NVIDIA Technologies
The global robotaxi market - physical AI's first commercial breakthrough - is projected to reach $400 billion by 2035, with over 6 million commercial.
Bypassing Inference Bottlenecks: Accelerating Complex AI Search with Retrieve-for-Train
Algorithms & Theory.
Accessing OpenAI Models on Amazon Bedrock From Australia with Global Cross-Region Inference
Australian teams can now access OpenAI GPT-5.6 Sol, Terra, and Luna models on Amazon Bedrock with global cross-Region inference from the Asia Pacific.
Inference
Inference is the process of using a trained AI model to make predictions or generate text based on new inputs. During inference, data flows forward through the neural network to produce an output, without modifying the model's weights.
Blackwell
Blackwell is NVIDIA's high-performance GPU architecture designed specifically to accelerate trillion-parameter large language models, offering massive throughput improvements for AI training and inference workloads.
NVIDIA
NVIDIA is a pioneer of GPU computing, dominating the hardware market for AI acceleration, training, and inference with its high-performance Hopper and Blackwell architectures.
Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:
Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.