Most data center servers fit one, two, four, or eight GPUs, with the specific number determined by physical PCIe slot count, power delivery capacity, and whether the system uses PCIe cards or an SXM baseboard. Single and dual GPU configurations are common in standard 1U or 2U rack servers using PCIe cards such as the L40S, RTX PRO 6000, or H100 PCIe, offering flexibility and lower power requirements for smaller workloads. Four GPU configurations occupy a middle ground, often used for departmental inference or moderate fine tuning workloads. Eight GPU configurations, built around NVIDIA's HGX or DGX baseboard with SXM modules and full NVLink connectivity, are the standard for serious training and high concurrency inference, and represent the densest common configuration since going beyond eight GPUs within a single server generally runs into practical limits around power delivery, cooling, and PCIe or NVLink topology complexity. Beyond eight GPUs per node, scaling to more compute means connecting multiple 8 GPU servers together over InfiniBand or high speed Ethernet rather than cramming more GPUs into one chassis, as seen in GB200 NVL72 rack scale systems that link many nodes into one larger NVLink domain. Nanobase AI recommends GPU count per server based on parallelism needs rather than simply maximizing density.

Slot count, power, and topology set the ceiling, not ambition

The number of GPUs a single server can hold is determined by three physical constraints working together: how many PCIe slots or SXM baseboard sockets the chassis has, how much power the system's power supplies and distribution can deliver, and how GPU-to-GPU communication (PCIe versus NVLink) is topologically wired. Most data center servers land on one, two, four, or eight GPUs specifically because these are the configurations where slot count, power delivery, and interconnect topology all work out cleanly, not because other numbers are technically impossible.

Choosing among these configurations is a workload-sizing decision, not simply a matter of buying the biggest box available, since a larger GPU count per server adds cost, power, and cooling complexity that only pays off when the workload can actually use tighter, faster inter-GPU communication.

Common configurations at a glance

GPU countTypical form factorInterconnectCommon use case
1–21U–2U rack server, PCIe cards (e.g., L40S, RTX PRO 6000, H100 PCIe)PCIe, occasionally NVLink bridge for 2 GPUsSmall inference workloads, departmental serving, dev/test
42U–4U rack server, PCIe or partial NVLinkPCIe or partial NVLink meshModerate inference or fine-tuning, mid-size deployments
8HGX or DGX baseboard, SXM modulesFull NVLink mesh across all 8 GPUsSerious training, high-concurrency inference, multi-GPU model parallelism
Beyond 8Multiple 8-GPU nodes linked over networkInfiniBand or high-speed Ethernet between nodesLarge-scale training, rack-scale systems like GB200 NVL72

Why PCIe configurations top out lower than SXM ones

PCIe-based servers connect GPUs through standard expansion slots, which are practical to fit in numbers of one, two, or four within typical 1U to 4U chassis without exotic power delivery, and communication between PCIe GPUs, absent a direct NVLink bridge, has to route through the CPU or PCIe switch fabric, adding latency compared to a dedicated GPU-to-GPU link. SXM-based systems use a specialized baseboard, standardized by NVIDIA's HGX design, that provides both higher per-GPU power delivery and a full NVLink mesh connecting every GPU to every other GPU at high bandwidth, which is what makes the 8-GPU configuration the practical ceiling for a single chassis: going further would require either splitting the NVLink domain or adding a second baseboard, both of which push the design toward a multi-node architecture instead.

When to scale within a server versus across servers

  1. If a model and its KV cache fit comfortably within one to two GPUs' memory, a 1–2 GPU PCIe server is usually the simplest and most cost-effective choice.
  2. If the workload needs model parallelism across a handful of GPUs but does not require the tightest possible inter-GPU bandwidth, a 4 GPU configuration is a reasonable middle ground.
  3. If training or high-concurrency inference genuinely benefits from full NVLink bandwidth across every GPU in the job, an 8 GPU HGX or DGX-class server is the standard choice.
  4. If compute needs exceed what a single 8 GPU node can provide, the next step is connecting multiple 8 GPU nodes over InfiniBand or high-speed Ethernet rather than seeking a denser single-chassis configuration, which is how rack-scale systems like GB200 NVL72 are architected.

Frequently asked questions

Is an 8 GPU server always better than a 4 GPU server?

Not for every workload. An 8 GPU server offers more compute and full NVLink bandwidth across all GPUs, but it also costs more, draws more power, and requires more cooling capacity, so it is only the better choice when the workload can genuinely use that scale.

Can PCIe GPUs communicate with each other without going through the CPU?

Some PCIe GPUs support direct NVLink bridges between pairs, bypassing the CPU for that specific link, but this typically connects only two GPUs at a time rather than providing the full mesh that an SXM-based 8 GPU system offers.

Why do most 8 GPU servers use SXM instead of PCIe modules?

SXM modules connect to a dedicated baseboard that provides higher power delivery per GPU and a full NVLink mesh across all eight, which PCIe slot-based designs cannot replicate at that density without significantly more complex bridging.

What comes after an 8 GPU server if we need more compute?

The standard path is connecting multiple 8 GPU nodes together over InfiniBand or high-speed Ethernet rather than trying to fit more GPUs into a single chassis, which is the architecture behind larger training clusters and rack-scale systems.

How Nanobase AI helps

Nanobase AI recommends GPU count per server based on actual parallelism and concurrency needs rather than simply maximizing density, sizing between 1, 2, 4, and 8 GPU configurations, and multi-node clusters where appropriate. This sizing work draws on our broader guidance on how many GPUs are needed for 70B and 405B class models. Explore GPU infrastructure services or get in touch to size a configuration for your specific model and traffic pattern.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.