A fully configured GB200 NVL72 rack draws approximately 120 kilowatts, a figure NVIDIA has cited for the complete rack containing 36 Grace CPUs and 72 B200 GPUs connected through a single NVLink domain, though actual draw varies somewhat with workload and specific configuration. That power density is roughly ten to twenty times higher than a traditional air cooled server rack, which is why the GB200 NVL72 is designed from the ground up around direct to chip liquid cooling rather than air, since no practical amount of airflow could remove that much heat from a single rack footprint. Supporting a rack at this density also requires reinforced electrical infrastructure, including high capacity power distribution units and often facility level upgrades to bring sufficient three phase power to the row, well beyond what most existing enterprise data centers or colocation suites provide without significant retrofitting. This is one of the main reasons GB200 NVL72 deployments are concentrated among hyperscalers and a small number of purpose built AI data centers rather than typical enterprise server rooms. Organizations considering this scale should engage facility engineers early in the planning process. Nanobase AI advises most enterprise clients toward lower density H100, H200, or B200 node deployments that fit within conventional data center power envelopes.

120 kW is a rack, not a server, and that distinction matters

The approximately 120 kilowatt figure NVIDIA cites for a fully configured GB200 NVL72 rack describes an entire rack containing 36 Grace CPUs and 72 B200 GPUs connected through a single NVLink domain, not an individual server. A single rack drawing roughly 120 kW is somewhere in the range of ten to twenty times the power density of a traditional air-cooled server rack, which historically ran anywhere from a few kilowatts up to about 15–20 kW at the high end. That difference in scale is why GB200 NVL72 planning cannot reuse assumptions from prior-generation rack deployments, even ones that already included H100 or B200 nodes.

Because the 72 GPUs sit inside one NVLink domain, the rack is also engineered as a single logical unit rather than 72 independent servers that happen to share a cabinet, which shapes both the physical layout and how power and cooling are distributed across it.

What 120 kW requires at the facility level

RequirementTypical prior-generation rack (H100-era)GB200 NVL72 rack
Power drawRoughly a few kW up to ~15–20 kWApproximately 120 kW
Cooling methodAir cooling with hot-aisle containment, commonly sufficientDirect-to-chip liquid cooling, effectively required
Electrical serviceStandard single or three-phase circuits per rackReinforced high-capacity three-phase distribution, often facility-level upgrades
Facility fitMost existing enterprise data centers and colocation suitesConcentrated among hyperscalers and purpose-built AI data centers
Deployment scopeIndividual racks added incrementallyRow- or hall-level planning due to cumulative density

No amount of increased airflow makes air cooling practical at 120 kW in a single rack footprint; the physics of heat removal at that density require the far higher thermal transfer efficiency of liquid, which is why GB200 NVL72 is designed from the ground up around direct-to-chip cooling and a coolant distribution unit rather than offering air cooling as an option.

Why most enterprises will not deploy this directly

Bringing a single rack to roughly 120 kW is not just a cooling problem; it requires facility-level electrical upgrades to deliver that much three-phase power to one row, which very few existing enterprise data centers or colocation suites have provisioned for, since typical facilities are designed around a much lower average watts-per-rack figure across the whole floor. Retrofitting an existing facility for even a handful of GB200 NVL72 racks can mean upgrading switchgear, transformers, and the chilled water loop feeding the row, work that is closer to a facility construction project than a rack installation.

This is the main reason GB200 NVL72 deployments are concentrated among hyperscalers and a small number of purpose-built AI data centers rather than showing up broadly in typical enterprise server rooms. Most enterprise buyers evaluating Blackwell-generation hardware are better served by lower-density HGX B200 nodes or H100/H200 systems that fit within conventional data center power envelopes, reserving rack-scale NVL72 deployments for organizations with access to purpose-built or hyperscale facilities.

Planning steps if rack-scale density is genuinely on the roadmap

  1. Engage facility and electrical engineers early, before any hardware procurement decision, since lead times for electrical infrastructure upgrades typically exceed hardware lead times.
  2. Confirm whether the target facility (owned or colocation) can deliver the required three-phase capacity per row, not just per rack.
  3. Plan the chilled water loop and coolant distribution unit sizing at the row or hall level rather than per rack, since NVL72 deployments are rarely a single isolated unit.
  4. Model whether a smaller number of lower-density HGX B200 or H100/H200 nodes meets the actual workload requirement before committing to rack-scale infrastructure.

Frequently asked questions

Does every GB200 NVL72 deployment draw exactly 120 kW?

Approximately 120 kW is the figure NVIDIA cites for a fully configured rack; actual draw varies somewhat with workload and specific configuration, so facility planning should use it as a design target with reasonable margin rather than an exact constant.

Can a colocation facility support a GB200 NVL72 rack?

Only a small number of purpose-built, high-density colocation facilities currently support this level of power and liquid cooling per rack. Most standard colocation suites are provisioned well below this density and would require significant facility upgrades.

Is GB200 NVL72 the right choice for a mid-size enterprise?

Rarely, for enterprises without access to purpose-built high-density facilities. Lower-density H100, H200, or standard HGX B200 deployments generally fit conventional enterprise data centers far more easily and meet most non-hyperscale workload requirements.

How does GB200 NVL72 relate to a standard HGX B200 node?

They use the same Blackwell GPU generation but are architecturally different: an HGX B200 node is an 8-GPU server, while GB200 NVL72 is a rack-scale system linking 72 GPUs and 36 Grace CPUs into one large NVLink domain, requiring a different order of facility infrastructure.

How Nanobase AI helps

Nanobase AI, a Silicon Valley enterprise AI engineering company, advises most enterprise clients toward right-sized H100, H200, or HGX B200 deployments that fit within conventional data center power envelopes, and only recommends rack-scale GB200 NVL72 planning when a client's facility access and workload genuinely justify it. Read our comparison of H100 vs H200 vs B200 for LLM inference or explore our GPU infrastructure solutions to size the right density for your facility.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.