
Cognition Becomes First Customer for NVIDIA Vera Rubin NVL72 on CoreWeave Cloud
AI Executive Summary
CoreWeave announced the limited availability of the NVIDIA Vera Rubin NVL72 platform on CoreWeave Cloud, integrating 72 Rubin GPU and 36 Vera CPUs.
Cognition became the first production customer, deploying Devin to run agentic AI workloads and benchmarking a 4.8x increase in total token throughput.
Why It Matters
Strategic TakeawayThe deployment of rack-scale liquid-cooled architectures featuring HBM4 memory bandwidth and NVLink connectivity fundamentally alters agentic AI unit economics by cutting token costs by 45x compared to Blackwell. This removes major computational bottlenecks for long-context reasoning and multi-step autonomous workflows.
Multi-Vector Implications
- TECHNICALIntegration of 72 Rubin GPU and 36 Vera CPUs via NVLink 6 delivers 1,400 TB/s HBM4 bandwidth, slashing multi-step agentic inference latency.
- MARKETCoreWeave captures high-value autonomous agent customers like Cognition early, establishing a competitive moat in next-gen hardware deployment.
- GOVERNANCERapid bare-metal rack handovers require standardized multi-region security and compliance automation across managed Kubernetes and object storage.
Strategic Outlook
12-18M HorizonOver the next 12-18 months, cloud providers will aggressively race to general availability for Vera Rubin architectures to support enterprise scaling of trillion-parameter autonomous agent. Hardware density and liquid-cooling integration will become the primary competitive differentiators for specialized AI clouds.
Referenced Coverage & Sources
Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference V6.1 Debut
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics.
Right-size Generative AI Endpoints with Concurrency Sweeps on Amazon SageMaker AI
Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels.
Kimi K3: a Claude Clone or Something Else?
Take a closer look at Kimi K3, its architecture, benchmark performance, and reported similarities to Claude, and examine what the evidence says about how distinct the model really is.
Tutorial: Benchmarking GPT-6 Astra Vs Claude Fable 5.1 Vs GPT-5.6 Sol Using W&B Weave
Learn how to benchmark GPT-6 Astra, Claude Fable 5.1, and GPT-5.6 Sol with Weave, comparing model quality, latency, cost, and performance across practical evaluation tasks.
Agentic AI
Agentic AI refers to artificial intelligence systems designed to act autonomously, make decisions, plan workflows, and execute tasks without constant human intervention. Unlike traditional models that only respond to queries, agentic systems use an agentic loop to perceive environments, reason over goals, use tools, and iterate to achieve outcomes.
NVIDIA
NVIDIA is a pioneer of GPU computing, dominating the hardware market for AI acceleration, training, and inference with its high-performance Hopper and Blackwell architectures.
Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:
Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.