CoreWeave has announced the first measured performance results for NVIDIA’s Vera Rubin NVL72, demonstrating a tenfold increase in tokens per megawatt compared to the prior Blackwell architecture. This performance enhancement is significant as the industry increasingly prioritizes tokens per megawatt as a vital metric for evaluating AI inference hardware efficiency, moving beyond traditional theoretical FLOPS. Additionally, major cloud providers are preparing to incorporate next-generation NVIDIA AI systems into their data center networks to meet the growing demands of AI applications.
NVIDIA: NVIDIA designs and manufactures advanced GPUs and full-stack AI computing platforms that power modern inference and training workloads. Its Vera Rubin NVL72 architecture introduces targeted improvements in memory, networking, and transformer-specific optimizations. Recent measured results from partners confirm its role in advancing power-efficient AI silicon.
CoreWeave: CoreWeave is an AI-focused cloud infrastructure provider that delivers specialized GPU clusters optimized for large-scale training and inference workloads. It has delivered the first real-world silicon measurements for NVIDIA’s Vera Rubin NVL72 system running on its live hardware. This positions CoreWeave as an early validator of next-generation AI accelerator performance in production environments.
Vera Rubin NVL72: Vera Rubin NVL72 is NVIDIA’s next-generation AI computing rack platform designed as the successor to Blackwell, with architecture enhancements for high-throughput inference. It incorporates HBM4 memory and advanced interconnects to improve real-world token generation efficiency. Early measured performance on DeepSeek-R1 has been validated on live systems by infrastructure partners.
Efficiency Metric: Industry discussions now emphasize tokens per megawatt as a practical benchmark for evaluating AI inference hardware beyond theoretical FLOPS.
Platform Deployment: Major cloud providers are advancing plans to integrate next-generation NVIDIA AI systems into global data center networks to support expanding AI factory requirements.
