An 8-GPU H100 server typically draws somewhere between about 8 and 11 kilowatts under sustained full load, with the GPUs themselves accounting for roughly 5.6 kilowatts at their 700 watt SXM5 rating and the rest coming from CPUs, memory, storage, and cooling fans. Running continuously at full load for a full year works out to roughly 70,000 to 95,000 kilowatt-hours of electricity, which at typical commercial electricity rates of about 0.08 to 0.20 dollars per kilowatt-hour translates into an annual electricity cost of very roughly 6,000 to 19,000 dollars for that single server, before accounting for data center cooling overhead. Facilities with inefficient cooling can add another 30 to 50 percent through the power usage effectiveness ratio, since every watt drawn also has to be removed as heat. Actual annual cost will be lower for workloads that do not run at sustained peak power, such as bursty inference with idle periods, and higher for continuous training jobs. Multiplying this per-server figure by the number of nodes in a cluster is the fastest way to estimate total facility electricity spend for budgeting purposes. Local commercial electricity rates as of 2026 should be used for an accurate number rather than a national average. Nanobase AI, an NVIDIA Inception Program member, includes detailed power and cooling cost estimates in every GPU infrastructure proposal.

The formula is simple, the inputs are where estimates go wrong

Annual electricity cost for an H100 server follows a straightforward formula: power draw in kilowatts, times hours running, times the local electricity rate, times a facility overhead factor for cooling. The formula itself is not where budgets go wrong; the two inputs most often estimated poorly are the facility overhead factor and the assumption that the server runs at full sustained load around the clock, and both can move the final number by 30% or more in either direction.

Building the formula step by step

StepInputTypical range
1. GPU power draw8x H100 SXM5 at 700W rated eachRoughly 5.6 kW from GPUs alone
2. Full-server power drawGPUs plus CPUs, RAM, storage, fansRoughly 8-11 kW under sustained full load
3. Annual energy at full loadPower draw x 8,760 hoursRoughly 70,000-95,000 kWh per year
4. Electricity costAnnual energy x local commercial rateVerify current local commercial rate as of 2026
5. Facility overhead (PUE)Cooling and facility power on top of IT loadAdds roughly 30-50% in an inefficiently cooled site

The formula: Annual electricity cost = Power draw (kW) x Hours run x Electricity rate ($/kWh) x PUE multiplier.

Why load pattern changes the answer more than people expect

A server assumed to run at sustained full load around the clock produces the highest possible electricity estimate, but real workloads rarely hold that pattern; bursty inference with meaningful idle periods draws noticeably less average power than continuous training, and using the wrong assumption for the workload type is one of the most common ways electricity budgets miss the mark. A continuous training job is a reasonable fit for the full-load assumption; a customer-facing inference endpoint with daily traffic peaks and overnight lulls is not, and modeling it as if it were overstates the annual electricity line item.

Why PUE deserves its own line item

Power usage effectiveness measures how much total facility power is drawn for every watt that reaches the actual compute hardware, and it is the difference between a well-run modern data center and an older or poorly cooled space. A facility running an efficient PUE adds a modest overhead on top of the GPU's own draw, while an inefficiently cooled space can add 30-50% or more, since every watt the GPU draws also has to be removed as heat, and that removal itself consumes power. This is a facility characteristic, not a GPU characteristic, which is why the identical server can carry different annual electricity costs depending purely on where it is housed.

Scaling the formula to a cluster

Multiplying the single-server annual figure by node count gives a first-pass cluster estimate, but multi-node clusters also add networking equipment power draw, which is typically small per node but not zero, and should be included for larger clusters rather than assumed negligible. Redundant cooling and power infrastructure at the facility level can also shift the effective PUE for a dedicated GPU cluster compared with a mixed-workload data center.

Frequently asked questions

Does the H200 or B200 change this electricity formula meaningfully?

The formula stays the same, but the power draw input changes: H200 shares the H100's roughly 700W per-GPU rating, while B200-generation GPUs draw more per card, so the same formula with an updated power draw figure applies across generations.

How much does idle time reduce the annual electricity estimate?

It depends on the actual idle-to-load ratio of the workload; a server idling at a fraction of peak power for a meaningful share of the day can see materially lower annual energy use than the full-load assumption, so measuring actual load pattern gives a more accurate number than assuming continuous peak draw.

Should electricity cost be compared against cloud rental rates directly?

Electricity is only one component of on-prem total cost of ownership; a fair comparison against cloud rental rates needs the full TCO figure, including amortized hardware, support, and staff time, not electricity cost in isolation.

Where can a realistic local commercial electricity rate be found?

Local utility rate schedules or a recent facility electricity bill are the most reliable sources, since commercial electricity rates vary meaningfully by region and by contract terms, and a national average can be off by a wide margin for a specific location.

How Nanobase AI helps

Nanobase AI, an NVIDIA Inception Program member, includes detailed power and cooling cost estimates in every GPU infrastructure proposal, using the target facility's actual PUE and the workload's real load pattern rather than a full-load assumption. This feeds directly into the full on-prem TCO model and colocation cost planning.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.