NVIDIA has announced significant advancements in its full-stack inference software for its Blackwell GPUs, highlighting a remarkable 5x reduction in token cost through software optimizations in just one month, even using the same hardware. This improvement is crucial as organizations transition from AI pilots to production, where efficiency metrics are shifting towards cost per token and performance. The software is designed to leverage continuous gains from a stack of optimizations that work in unison, enabling leading companies to enhance tasks across various domains, including healthcare and coding, by increasing throughput significantly. Major open-source frameworks, like PyTorch, are also evolving alongside NVIDIA’s technology, ensuring that every new AI breakthrough is immediately integrated into its ecosystem.
NVIDIA: NVIDIA develops graphics processing units and full-stack AI infrastructure solutions. The news centers on its inference software optimizations that compound performance gains on Blackwell hardware through techniques like disaggregated serving and NVFP4 precision. Open source frameworks such as PyTorch integrate tightly with NVIDIA architectures to accelerate new AI models from the start.
Baseten: Baseten provides a platform for deploying and serving machine learning models. It is cited in the news as an early adopter that applied TensorRT-LLM with proprietary optimizations on NVIDIA Blackwell to improve token throughput for reasoning and long-context tasks.
PyTorch: PyTorch is an open source machine learning framework widely used for AI research and production. The news emphasizes its deep integration with NVIDIA CUDA, allowing every new AI advancement to run natively on NVIDIA hardware from launch.
Cognition: Cognition builds AI systems focused on reinforcement learning and agentic workflows. The news highlights its use of NVIDIA Dynamo to manage inference GPUs on Blackwell, enabling scalable handling of inference engine lifecycles.
Deep Infra: Deep Infra operates a cloud platform for running open source AI models. It is noted for deploying frontier models on NVIDIA Blackwell from day one using the full inference software stack, including support for models like DeepSeek V4.
Together AI: Together AI offers a platform for training and inference of open source models. It is mentioned for applying TensorRT-LLM on Blackwell to help Cursor move from model optimizations to reliable production endpoints for real-time coding applications.
DigitalOcean: DigitalOcean delivers cloud infrastructure and managed services for developers and businesses. The news features its work enabling Hippocratic AI to leverage NVIDIA inference software on Blackwell, achieving higher throughput while meeting strict latency requirements for production workloads.
Production Adoption: Leading AI companies are already using NVIDIA’s full-stack inference software on Blackwell to scale workloads across reasoning, coding, and specialized domains such as healthcare.
Software Optimization: NVIDIA’s full inference stack enables continuous performance gains on deployed hardware through layered improvements that compound when used together.
Open Source Integration: Major open source frameworks co-evolve with NVIDIA architectures, ensuring new AI breakthroughs run on the company’s platforms immediately upon release.
