There is no single fixed price for an H100 server because the total depends on GPU form factor, GPU count, and everything built around the GPUs. An 8-GPU H100 SXM5 node with dual server-class CPUs, 2 terabytes or more of system RAM, several NVMe drives, InfiniBand or high-speed Ethernet networking, and a multi-year support contract has historically been quoted by system integrators in a broad range of roughly 200,000 to 350,000 dollars, while smaller PCIe-based configurations with one to four cards cost a small fraction of that. Networking cards, redundant power supplies, and enterprise support and warranty terms typically add 10 to 20 percent on top of the base hardware bill of materials. Pricing has also trended down as H200 and Blackwell generation GPUs pull demand away from Hopper, so quotes vary meaningfully between vendors and shift with GPU supply. As of 2026, buyers should treat any number as a starting point and get a current, itemized quote rather than relying on list prices found online. Nanobase AI, a Silicon Valley enterprise AI engineering company, sources and configures H100 servers from qualified integrators and provides itemized, fixed-scope quotes based on a customer's actual workload.

Five line items make up the price, not one

Quoting "an H100 server" as a single figure hides the fact that the number is really a sum of five independent line items: the GPU modules themselves, the host system (CPUs, RAM, motherboard, chassis, power supplies), storage, networking, and support or warranty terms. Two configurations both labeled "H100 server" can differ by a wide margin once GPU count, form factor, and support tier are held apart from the base chassis cost. A single-GPU PCIe workstation and an 8-GPU SXM5 node with InfiniBand share almost nothing in common except the GPU brand.

GPU count and form factor are the biggest lever

The single largest driver of total price is how many GPUs the node carries and whether they are SXM5 (the module used in 8-GPU NVLink-connected nodes) or PCIe (used in 1-4 GPU add-in-card configurations). SXM5 nodes cost more per GPU because they require NVSwitch fabric, denser power delivery, and liquid or high-static-pressure air cooling that PCIe cards do not need. The GPU line item itself typically represents the majority of the bill of materials, but it is rarely the only thing driving quote-to-quote variance.

Cost componentWhat it coversRelative weight in total price
GPU modulesThe H100 SXM5 or PCIe cards themselvesLargest single share
Host platformDual server CPUs, 1-2 TB+ RAM, motherboard, chassis, PSUsMeaningful, scales slowly with GPU count
StorageNVMe drives for model weights, checkpoints, cacheSmall, but grows with dataset and checkpoint size
NetworkingInfiniBand or 400G Ethernet NICs and switchesModerate, rises sharply for multi-node clusters
Support and servicesMulti-year warranty, onsite support, firmware validationTypically 10-20% on top of hardware

What "support and services" actually buys

Enterprise buyers often treat the support line as a discretionary add-on, but for GPU hardware it functions closer to insurance against a very expensive failure mode: a dead GPU or NVSwitch mid-training run. The support contract is usually the cheapest line item on the quote and the most expensive one to skip if something fails at 2 a.m. during a production inference workload. Multi-year next-business-day or 4-hour onsite terms typically add a percentage on top of hardware cost rather than a flat fee, so the add-on scales with the size of the deployment being protected.

A worked structure for budgeting a quote

Rather than anchoring on a headline number found online, a more reliable way to sanity-check an incoming quote is to build it up from its parts:

  1. Price the GPU line item for the exact SKU and count (SXM5 8-GPU node versus PCIe 1-4 GPU card).
  2. Add the host platform cost, which stays relatively flat whether the node holds 4 or 8 GPUs.
  3. Add storage sized to the model weights and checkpoint volume the workload actually needs.
  4. Add networking, which jumps meaningfully once more than one node needs to talk to another over InfiniBand.
  5. Apply a 10-20% support and services markup on the hardware subtotal to get the all-in quote.
  6. Compare the result against at least two integrator quotes, since GPU allocation and vendor margin vary by supplier as of 2026.

Why quotes move independently of the GPU itself

GPU allocation constraints, not manufacturing cost, have been the dominant swing factor in H100 pricing since the chip's release, and that dynamic has not disappeared even as H200 and Blackwell-generation GPUs pull demand away from Hopper. A vendor with spare allocation may quote meaningfully below one that is capacity-constrained, independent of any difference in the underlying hardware. As of 2026, always request a current, itemized quote and verify current pricing rather than relying on a number seen in an older article or forum post.

Frequently asked questions

Does buying more GPUs per node reduce the per-GPU price?

Somewhat. The host platform, chassis, and networking costs are largely fixed regardless of whether a node holds 4 or 8 GPUs, so spreading them across more GPUs lowers the effective per-GPU overhead, though the GPU modules themselves are priced per unit regardless of node size.

Is PCIe meaningfully cheaper than SXM5 for the same GPU count?

Yes, PCIe configurations skip the NVSwitch fabric and denser cooling that SXM5 nodes require, which lowers both the GPU and host platform cost, at the tradeoff of lower GPU-to-GPU bandwidth that matters for multi-GPU model parallelism.

Should we buy directly from NVIDIA or through an integrator?

Most enterprise buyers purchase through a systems integrator or reseller rather than NVIDIA directly, since integrators assemble, validate, and support the full node and can often secure GPU allocation that individual buyers cannot access on their own.

How much does networking add for a multi-node cluster?

Networking cost rises sharply once GPUs need to communicate across nodes rather than within one, since InfiniBand switches, cables, and NICs scale with node count; a single 8-GPU node needs far less networking investment than a multi-node training cluster.

How Nanobase AI helps

Nanobase AI sources and configures H100 servers from qualified integrators and breaks every quote into its underlying line items so customers see exactly what they are paying for rather than a single opaque number. This is useful groundwork before comparing H100, H200, and B200 for inference or working through how many GPUs a target model actually needs. See /solutions for how sizing and procurement fit together, or view a live example at /demo.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.