Planning VRAM for the next three years is inherently uncertain, since both frontier model sizes and quantization efficiency have continued to shift, but the safer approach is to buy meaningfully more memory headroom than the current model requires rather than sizing to exactly what today's chosen model needs. GPUs like the H200, with 141 GB of HBM3e, or the B200, with roughly 180 GB, give substantially more room to adopt larger or higher-quality models later than an H100 at 80 GB does, and that headroom tends to matter more over a three-year horizon than raw compute speed, since new quantization formats like NVFP4 continue to make model deployment more memory-efficient even as top-end model sizes grow. A reasonable planning rule is to budget for roughly double the current model's memory footprint, giving room for either a larger model at the same quantization level or the same model at higher precision, without assuming any specific future model release. Software and interconnect also age less gracefully than raw memory capacity, so prioritizing GPUs with strong NVLink or InfiniBand support protects the ability to scale out later. Nanobase AI, a Silicon Valley enterprise AI engineering company, builds this multi-year headroom into every GPU cluster it designs rather than sizing to only the first workload.
Why memory headroom ages better than compute speed
Over a multi-year hardware planning horizon, two trends move in ways that are hard to predict individually but combine to favor buying memory headroom over buying raw speed. Model quality per parameter tends to improve over time, letting smaller models do more, while quantization formats keep getting more efficient, letting the same model fit in less memory. Both trends reduce future memory need per unit of capability. But new frontier models and new use cases (longer context, multimodal, agentic workflows with more concurrent tool calls) tend to raise the ceiling on what a leading deployment wants to run. The safer bet is not predicting which of these forces wins, but buying enough memory headroom to absorb either outcome without a forced early hardware refresh.
A generational memory comparison for planning purposes
| GPU generation | Memory | Memory bandwidth | Headroom over an 80 GB baseline |
|---|---|---|---|
| H100 | 80 GB HBM3 | 3.35 TB/s | Baseline |
| H200 | 141 GB HBM3e | 4.8 TB/s | ~76% more memory |
| B200 | ~180 GB HBM3e | ~8 TB/s | ~125% more memory, roughly double the bandwidth |
Bandwidth matters for throughput today, but for a three-year planning horizon specifically, the memory capacity column is what determines whether a larger or higher-precision future model can be adopted without a hardware change, which is why memory headroom, not bandwidth, should weight the decision most heavily for long-horizon purchases.
A concrete planning rule
A reasonable planning rule is to budget for roughly double the current model's memory footprint, giving room for either a larger model at the same quantization level or the same model at higher precision, without assuming any specific future model release. Concretely: if today's chosen model needs 70 GB at FP8 on an H100 with only about 10 GB of headroom, that configuration has essentially no room to grow; the same 70 GB model on an H200 (141 GB) leaves roughly double the model's own footprint in reserve, enough to move to a larger model, a higher precision, or meaningfully more concurrency and context length without new hardware.
- Calculate today's model footprint at the intended production precision, using the standard weight-size formula.
- Target total GPU memory at roughly double that footprint, not just enough to fit today's model with a thin margin.
- Weight GPUs with strong NVLink or InfiniBand support more heavily than raw clock speed differences, since interconnect quality protects the ability to scale out to more GPUs later, an option that ages less gracefully than memory capacity if the wrong interconnect generation is chosen.
- Revisit the plan annually rather than committing to a fixed three-year roadmap upfront, since quantization formats like NVFP4 continue to shift the memory-versus-quality trade-off in ways that are difficult to predict precisely years in advance.
Why software and interconnect choices matter as much as the memory number
Software and interconnect also age less gracefully than raw memory capacity: a GPU generation with weak driver or framework support for emerging formats, or a server chassis without NVLink bridges installed, limits future flexibility even if the memory number on paper looks generous. Prioritizing GPUs with strong ecosystem support and interconnect options protects the ability to scale out later, which matters as much for a three-year plan as the raw gigabyte figure does.
Frequently asked questions
Should I buy H200 or B200 if I'm planning three years out?
Either provides substantially more headroom than H100; the choice between them typically comes down to current availability, ecosystem maturity for the specific serving stack being used, and budget, all of which should be verified as of 2026 rather than assumed from older comparisons, as discussed in H100 or B200 for a 70B deployment.
Does buying more memory now actually save money over buying twice?
Often, yes, since a mid-life hardware upgrade carries its own disruption cost, migration effort, and potential downtime beyond the raw GPU price, making a larger upfront memory purchase frequently cheaper in total cost of ownership than two smaller purchases over the same period.
Will future quantization improvements make today's memory headroom unnecessary?
Possibly for a given model size, but new use cases and larger models tend to expand to fill available headroom over time, so betting entirely on quantization efficiency to avoid buying memory headroom is a riskier assumption than budgeting for both trends together.
Is doubling the memory footprint the right rule for every organization?
It is a reasonable general starting point, not a universal rule; organizations with a very stable, well-defined single use case and low growth expectations can reasonably plan closer to their current footprint, while those expecting to expand use cases should lean toward more headroom than the doubling rule suggests.
How Nanobase AI helps
Nanobase AI, a Silicon Valley enterprise AI engineering company, builds this multi-year headroom into every GPU cluster it designs rather than sizing to only the first workload, weighing memory capacity, interconnect and software ecosystem together for a purchase that holds up over the planning horizon.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.