NAVIGATION
Futuristic green Nvidia brand banner with circuit boards and GPU architecture.
Product Launch
Source:CoreWeave

CoreWeave Leads Cloud Providers in MLPerf® Inference V6.1 Performance with NVIDIA Blackwell Ultra

40s Read

AI Executive Summary

CoreWeave achieved leading cloud provider inference throughput in MLPerf Inference v6.1 using NVIDIA Blackwell and Blackwell Ultra platforms.

Operating within the Datacenter Closed division's Available category across four NVIDIA platforms and four model families, CoreWeave utilized an NVIDIA GB300 NVL72 rack running the Qwen3-VL-235B-A22B model to sustain 1,196 queries per second in the server scenario and 1,135 samples per second in the offline scenario.

Why It Matters

Strategic Takeaway

Scaling multimodal and mixture-of-experts reasoning model requires tightly co-designed rack-scale hardware and serving software to compress inference costs and overcome memory bandwidth bottlenecks.

Multi-Vector Implications

  • TECHNICALContinuous optimization of serving stacks, scheduling algorithm, and operations unlocks higher throughput from existing NVIDIA GB300 NVL72 hardware without waiting for next-gen silicon.
  • MARKETHigh per-GPU throughput metrics in standardized benchmarks like MLPerf v6.1 give cloud providers a direct competitive advantage in lowering customer infrastructure costs for frontier-scale models.
  • GOVERNANCEStandardized MLPerf evaluations provide transparent, auditable performance baselines that enterprise buyers must mandate when validating cloud AI infrastructure SLAs for production workloads.

Strategic Outlook

12-18M Horizon

Over the next 12 to 18 months, cloud AI infrastructure competition will increasingly hinge on rack-scale software co-design and Day-1 availability of Blackwell Ultra systems to efficiently serve multimodal and MoE reasoning model at scale.

Referenced Coverage & Sources

Full Story Intelligence
High Signal Density

Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.

CoreWeave Leads Cloud Providers in MLPerf® Inference v6.1 Performance with NVIDIA Blackwell Ultra
CoreWeaveSep 16, 2026
Advertisement
Related Timeline Breakthroughs
View Full Live Feed →
Technical & Market Glossary Definitions
View Full Glossary →
AI ConceptModel Operations

Inference

Inference is the process of using a trained AI model to make predictions or generate text based on new inputs. During inference, data flows forward through the neural network to produce an output, without modifying the model's weights.

AI ConceptHardware & Infrastructure

Blackwell

Blackwell is NVIDIA's high-performance GPU architecture designed specifically to accelerate trillion-parameter large language models, offering massive throughput improvements for AI training and inference workloads.

AI ConceptHardware & Infrastructure

NVIDIA

NVIDIA is a pioneer of GPU computing, dominating the hardware market for AI acceleration, training, and inference with its high-performance Hopper and Blackwell architectures.

Frequently Asked Questions & Summary Briefing
See how CoreWeave's full-stack optimizations delivered leading inference throughput for NVIDIA Blackwell and Blackwell Ultra. Reported by CoreWeave, this update represents a key development in the Enterprise Product Launch category.
SPIDITS Intelligence Ecosystem

Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:

💬 Want real-time AI updates? Join our Discord server.

Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.

Join SPIDITS Discord →
CoreWeave Leads Cloud Providers in MLPerf® Inference V6.1 Performance with NVIDIA Blackwell Ultra | AI Timeline | SPIDITS AI