Yes, B200 and GB200 GPUs are available to rent in the cloud as of 2026, though availability and lead time vary considerably by provider and region given continued high demand for the newest NVIDIA Blackwell hardware. AWS offers B200 through P6-B200 instances and GB200 NVL72 capacity through P6e instances, Microsoft Azure provides GB200 based ND series virtual machines, and Google Cloud offers B200 through its A4 machine type and GB200 NVL72 with Grace CPUs through A4X, while neoclouds such as CoreWeave, Nebius, and Lambda have also brought up substantial GB200 NVL72 capacity, often with shorter waitlists than the largest hyperscalers during peak demand periods. GB200 NVL72 in particular represents a significant architectural step up, connecting 72 GPUs in a single NVLink domain for very large model training and inference, and generally requires committing to a meaningful capacity block rather than small on-demand allocations. Pricing for B200 and GB200 capacity is still settling as of 2026 and varies by commitment length and provider, so current rates should be confirmed directly rather than assumed from earlier generation pricing. Nanobase AI tracks B200 and GB200 availability across hyperscalers and neoclouds to help enterprises secure capacity for the largest model workloads.
Renting the hardware is the easy part; using it well is not
B200 and GB200 GPUs are broadly available to rent across AWS, Azure, Google Cloud, and established neoclouds as of 2026, so the practical question has shifted from "can we get access" to "is our software stack ready to actually use this hardware effectively." GB200 NVL72's rack-scale design, connecting 72 GPUs in a single NVLink domain, only delivers its advantage over standard multi-node clusters when the serving or training framework understands and exploits that topology, which is a software readiness question, not a procurement question. Renting the hardware without that readiness produces a cluster that runs, but at a fraction of its potential throughput.
B200 versus GB200 NVL72: two different commitments
| Aspect | B200 (standard instance) | GB200 NVL72 (rack-scale) |
|---|---|---|
| Typical footprint | Single instance, several GPUs | Full rack, 72 GPUs in one NVLink domain |
| Software readiness needed | Standard multi-GPU serving, similar to H100/H200 | Framework support for the larger NVLink domain and Grace CPU pairing |
| Minimum commitment | Often available in smaller on-demand allocations | Typically requires a meaningful capacity block commitment |
| Best fit | Standard large-model inference and training | Very large models needing the largest single coherent memory and bandwidth domain |
A team evaluating GB200 NVL72 should confirm that its serving framework, whether TensorRT-LLM or vLLM, has mature support for the topology before committing to a capacity block, since framework support for new hardware generations typically lags initial hardware availability by some months.
Where to actually get it
AWS offers B200 through P6-B200 instances and GB200 NVL72 capacity through P6e instances, Microsoft Azure provides GB200-based ND series virtual machines, and Google Cloud offers B200 through its A4 machine type with GB200 NVL72 available through A4X pairing GB200 with Grace CPUs. Neoclouds such as CoreWeave, Nebius, and Lambda have also brought up substantial GB200 NVL72 capacity, often with shorter waitlists than the largest hyperscalers during peak demand periods, though with a narrower set of adjacent enterprise services. Availability and lead time vary considerably by provider and region given continued high demand for the newest Blackwell hardware, so current allocation timelines should be confirmed directly.
What to check before committing to a capacity block
- Confirm the serving or training framework version in use has validated support for the specific GPU and topology, not just the GPU family in general.
- Size the actual workload against GB200 NVL72's advantage; a model that does not need the full 72-GPU coherent domain may see little benefit over a well-configured multi-node H200 cluster.
- Review the minimum commitment terms, since GB200 NVL72 capacity typically requires committing to a meaningful block rather than small on-demand allocations.
- Plan a validation period on a smaller allocation before committing to the full capacity block, if the provider allows it.
Skipping this checklist is how a well-funded team ends up with a fully booked GB200 NVL72 rack running at a fraction of its throughput because the serving stack was not ready for it.
Frequently asked questions
Is GB200 NVL72 always better than a cluster of separate B200 or H200 instances?
Not for every workload; GB200 NVL72's advantage comes from its 72-GPU single NVLink domain, which benefits models and training jobs that need very large coherent memory and bandwidth, but smaller models may not see a proportional benefit over a well-tuned standard multi-node cluster.
How much lead time should we expect to secure GB200 NVL72 capacity?
Lead times vary by provider and current demand, with neoclouds sometimes offering shorter waitlists than the largest hyperscalers during peak periods, so current availability should be checked directly with prospective providers rather than assumed from general market commentary or older announcements.
Do we need Grace CPUs to use GB200 effectively?
Grace CPU pairing, as offered in configurations like Google Cloud's A4X, provides tighter CPU-GPU memory coherence that benefits specific workloads, but standard GB200 configurations without Grace pairing are also available and sufficient for many inference and training use cases in practice.
Can we rent B200 without committing to a large capacity block?
Yes, standard B200 instances are often available in smaller on-demand allocations similar to previous GPU generations, while GB200 NVL72's rack-scale nature more often requires a meaningful capacity commitment given the scale of the deployment and its shared NVLink domain across the rack.
How Nanobase AI helps
Nanobase AI, an NVIDIA Inception Program member, tracks B200 and GB200 availability across hyperscalers and neoclouds and validates serving framework readiness before helping enterprises commit to a capacity block for the largest model workloads. This connects to what a GPU capacity block or reservation actually involves and to GPU generation comparisons for LLM inference.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.