
Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment
AI Executive Summary
NVIDIA AI factory operators face a $60 million per megawatt capital expenditure baseline, which NVIDIA addresses through full-stack codesign across its hardware, networking, and CUDA-X software.
According to SemiAnalysis AgentX data, next-generation NVIDIA Vera Rubin NVL72 systems achieve over 30x higher throughput per megawatt and up to 45x lower cost per million token compared to GB300 NVL72 systems on the DeepSeek V4 Pro model.
Despite rapid generational leaps, older hardware maintains long-term economic viability, demonstrated by the 2020-shipped NVIDIA A100 GPU remaining in active commercial service with CoreWeave bookings extended through 2029.
Why It Matters
Strategic TakeawayMaximizing token per second per megawatt dictates AI factory profitability under strict power constraints, transforming infrastructure investment from short-term capital deployment into multi-year asset lifecycles. Full-stack codesign mitigates the obsolescence risk of older hardware generations by sustaining demand across diverse, tiered workload complexities.
Multi-Vector Implications
- TECHNICALNVIDIA Vera Rubin NVL72 and GB300 NVL72 architectures utilize full-stack codesign across compute, networking, and CUDA-X libraries to drive 30x throughput-per-megawatt gains.
- MARKETOperators extend server depreciation schedules and maintain multi-year bookings for older silicon, such as CoreWeave booking 2020-era A100 GPU through 2029.
- GOVERNANCEAI factory operators enforce strict capital allocation models calibrated around a $60 million per megawatt baseline and projected token-per-second-per-megawatt metrics.
Strategic Outlook
12-18M HorizonOver the next 12 to 18 months, AI factory operators will increasingly prioritize token-per-second-per-megawatt efficiency metrics to justify $60 million per megawatt capital expenditure amid power constraint ceilings. The deployment of NVIDIA Vera Rubin NVL72 systems will accelerate high-density processing while secondary market continue to monetize older generations like H100 and A100 chips for less demanding workloads.
Referenced Coverage & Sources
Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.
Nvidia Ties AI Factory Economics to Tokens and Power Efficiency
Artificial intelligence factory economics increasingly depend on more than access to high-performance graphics processing units. As agentic systems draw on multiple models, databases and tools, the entire data center must work as one computing system.
AMD Acquires World Labs AI Startup, Upping the Ante Against Nvidia
The deal, which is expected to close by year's end, is worth $8.2 billion.
At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia
NVIDIA AI Day Singapore, which takes place Sept.
Cognition Becomes First Customer for NVIDIA Vera Rubin NVL72 on CoreWeave Cloud
NVIDIA Vera Rubin NVL72 is now in limited availability on CoreWeave Cloud. Cognition becomes the first customer running production agentic AI workloads and benchmarking performance.
NVIDIA
NVIDIA is a pioneer of GPU computing, dominating the hardware market for AI acceleration, training, and inference with its high-performance Hopper and Blackwell architectures.
Agentic AI
Agentic AI refers to artificial intelligence systems designed to act autonomously, make decisions, plan workflows, and execute tasks without constant human intervention. Unlike traditional models that only respond to queries, agentic systems use an agentic loop to perceive environments, reason over goals, use tools, and iterate to achieve outcomes.
Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:
Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.