
Photo: jurvetson (BY)
Hardware & ComputeNVIDIA’s Annual GPU Cadence and What It Means for Research Budgets
NVIDIA now ships a new data centre architecture every year. Blackwell arrived in 2025, Blackwell Ultra followed at the end of it, Rubin lands in 2026, Rubin Ultra is scheduled for 2027, and Feynman is already on the published roadmap for 2028. That cadence is a strategic fact with budget consequences, and university labs are the group least equipped to absorb it.
What actually changes with Rubin
The headline numbers are large. Rubin pairs HBM4 memory with a new generation of NVLink, and the rack-scale configurations are quoted at several times the dense low-precision throughput of the Blackwell generation they replace. Memory capacity per GPU has climbed to the point where models that previously required splitting across devices fit on one.
That last point matters more than the FLOPS figure. When a model fits in one device’s memory, you stop paying the communication cost of splitting it, and the effective speedup is larger than the raw compute comparison suggests. Conversely, when it does not fit, extra compute buys you less than the specification implies.
The problem an annual cadence creates for research
Enterprise buyers can depreciate hardware over three years and refresh continuously. A research lab typically buys from a grant, on a schedule set by the funding body rather than by the vendor, and then keeps the hardware for five to seven years.
Against an annual cadence, that means any purchase is two to three generations behind before it is retired. This is not a reason to wait — waiting is the one strategy guaranteed to fail, because there is always something better in twelve months. It is a reason to change what you optimise for.
Buy memory, not throughput
Compute improves fastest between generations. Memory capacity improves more slowly and constrains what you can run at all. A card with more VRAM and less compute will still be useful in four years, running workloads a faster card with less memory simply cannot hold. The reverse is not true.
Count the power and cooling before the card
The per-device power draw of current accelerators has outrun what a lot of university machine rooms were built for. It is entirely possible to win a grant, buy the hardware, and discover the building cannot power or cool it. Establish the facility constraint first; it determines the realistic option set more often than the budget does.
Consider the previous generation deliberately
When a new architecture ships, the prior generation’s price drops while its capability does not. For a lab whose workloads are bounded by memory and dataset size rather than by frontier-scale training, last year’s hardware at a discount is frequently the better purchase. This is unglamorous and it is usually correct.
Rent for peaks, own for baseline
The useful division is between what you run constantly and what you run occasionally. Development, iteration, and small experiments benefit from owned hardware: no queue, no hourly meter, no data movement. Occasional large training runs are better rented, because you get current hardware for the days you need it without owning it for the years you do not.
Labs that own enough to cover their peak load spend most of the year with expensive hardware sitting idle. Labs that rent everything discover that hourly billing punishes the exploratory work where most research time actually goes.
The honest summary
An annual release cadence means your hardware is always behind. Accepting that early leads to better decisions than trying to time the market. Buy for memory capacity and facility fit, buy the generation that is discounted rather than the one that is announced, and rent the peaks.
Get new posts by email
Occasional writing on post-quantum cryptography, blockchain security and digital forensics. No more than twice a month, and nothing else.


