
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
AI Executive Summary
Together's kernels team ported the open‑source ThunderKittens library to NVIDIA's Vera Rubin NVL72 GPU and rewrote the NVFP4 and FP8 GEMM kernels.
By exploiting new ISA feature—doubling operand bytes per step from 32 to 64—their GEMM reached over 22 PFLOPS, about 42% to 44% of the roofline previously, and now rivals cuBLAS and CuTe DSL performance.
Why It Matters
Strategic TakeawayThe work proves that targeted kernel redesign for a new GPU ISA can extract near‑peak compute throughput, closing the performance gap between open‑source and vendor‑optimized libraries.
Multi-Vector Implications
- TECHNICALDoubling K‑byte per step forces redesign of tile granularity and shared‑memory staging to keep tensor cores fed on Vera Rubin.
- MARKETMatching cuBLAS performance with an open‑source GEMM may attract cost‑sensitive AI compute providers to adopt ThunderKittens.
- GOVERNANCEOrganizations will need to audit the open‑source kernel’s licensing and security posture before deploying it in production clusters.
Strategic Outlook
12-18M HorizonOver the next 12‑18 months Together will likely push ThunderKittens beyond 30 PFLOPS on Vera Rubin, integrate the kernels into major AI frameworks, and extend the approach to upcoming NVIDIA architectures, driving broader open‑source adoption in high‑throughput AI workloads.
Referenced Coverage & Sources
Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference V6.1 Debut
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics.
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa.
Physical AI Takes the Wheel: How the World's Robotaxi Leaders Are Building with NVIDIA Technologies
The global robotaxi market - physical AI's first commercial breakthrough - is projected to reach $400 billion by 2035, with over 6 million commercial.
Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers
AI factories are the infrastructure of the intelligence era.
NVIDIA
NVIDIA is a pioneer of GPU computing, dominating the hardware market for AI acceleration, training, and inference with its high-performance Hopper and Blackwell architectures.
Agentic AI
Agentic AI refers to artificial intelligence systems designed to act autonomously, make decisions, plan workflows, and execute tasks without constant human intervention. Unlike traditional models that only respond to queries, agentic systems use an agentic loop to perceive environments, reason over goals, use tools, and iterate to achieve outcomes.
Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:
Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.