Nvidia's next-generation Vera Rubin GPU is generating buzz, promising a staggering 5x performance boost in inference and 3.5x in training compared to its predecessor, Blackwell. But before enterprises start salivating over Rubin, scheduled for release in the latter half of 2026, the current Blackwell architecture is proving it's far from obsolete. In fact, Blackwell is getting significantly faster right now.

Blackwell's Unexpected Growth Spurt

While the spotlight shines on Vera Rubin, Nvidia has quietly been working its magic on Blackwell. Just months after its release, Blackwell is showing remarkable performance gains thanks to a series of software optimizations. "We continue to optimize our inference and training stacks for the Blackwell architecture," says Dave Salvator, director of accelerated computing products at Nvidia.

These aren't minor tweaks; Nvidia claims Blackwell GPU performance has increased by up to 2.8x in inference in a mere three months. This impressive leap is attributed to innovations baked into the Nvidia TensorRT-LLM inference engine, and the best part? These improvements don't require any hardware upgrades.

The gains were measured using DeepSeek-R1, a massive 671-billion parameter mixture-of-experts model. Key technical innovations driving the performance boost include programmatic dependent launch (PDL), refined all-to-all communication, multi-token prediction (MTP), and the NVFP4 format. These advancements collectively reduce the cost per million tokens and allow existing infrastructure to handle higher request volumes with lower latency. This means cloud providers and enterprises can scale their AI services without immediately reaching for their wallets.

Training Gets a Shot in the Arm, Too

Blackwell isn't just about inference; it's also a workhorse for training large language models. Nvidia reports a 1.4x increase in training performance on the GB200 NVL72 system in just five months, all achieved without any hardware changes. The magic behind this boost lies in optimized training recipes and algorithmic refinements. Nvidia engineers have developed sophisticated training methods that effectively leverage NVFP4 precision, unlocking substantial performance from the existing silicon.

Blackwell Now, Rubin Later: A Phased Approach

So, what's the takeaway for enterprises navigating the rapidly evolving AI landscape? According to Salvator, while Vera Rubin promises game-changing performance and efficiency gains, Blackwell remains a market-leading platform for running state-of-the-art AI models. He notes that the Rubin is built to address the growing compute demands created by the increasing size of language models. The Rubin platform trains large MoE models in a quarter the number of GPUs, inference token generation with 10X more throughput per watt, and inference at 1/10th the cost per token compared to Blackwell, based on Nvidia's early testing results.

For organizations with existing Blackwell deployments, updating to the latest TensorRT-LLM versions offers immediate access to the 2.8x inference and 1.4x training improvements, translating to real cost savings. Those planning new deployments in the first half of 2026 should also consider Blackwell, as waiting for Rubin means delaying AI initiatives and potentially losing ground to competitors. However, if you are planning a major infrastructure overhaul for late 2026 or beyond, Vera Rubin should absolutely be on your radar, as it offers transformational economics for large-scale AI operations. The smartest strategy, it seems, is a phased deployment: leverage Blackwell for your immediate needs while strategically planning for Vera Rubin's arrival.

"The smartest strategy, it seems, is a phased deployment: leverage Blackwell for your immediate needs while strategically planning for Vera Rubin's arrival."

— Sarah Kim, Automatica Press

Nvidia's commitment to continuous optimization ensures that enterprises can maximize the value of their current investments without sacrificing long-term competitiveness. The Blackwell architecture continues to deliver tangible results, even as we eagerly await the arrival of Vera Rubin. The choice isn't about either/or, but rather about strategically leveraging both to stay ahead in the AI race.