How many GPUs fit in one server: PCIe lanes, power and airflow decide
- The CPU sets the first limit: 128 PCIe Gen5 lanes per AMD EPYC 9004/9005 socket, 96 per Intel Xeon 6900P. Each full-speed GPU wants sixteen of them
- Beyond four to six cards, vendors add PCIe switches: Dell XE7745 (8 × 600 W), HPE DL380a Gen12 (up to 10 double-width) and Supermicro 4U (10) all do; a few dual-EPYC designs put eight straight on the CPUs without a switch
- Power is the second limit: eight 600 W cards are 4.8 kW before CPUs and fans, and every vendor answers with four to eight 2,700–3,200 W supplies on 200–240 V
- Airflow is the third: about 155 CFM per kW at an 11 °C rise, so an eight-GPU box moves more air than a small room’s ventilation, and vendors cap inlet temperature at 30–35 °C
- The number that stops most orders is the feed: one 16 A single-phase circuit is 3.7 kW, and one eight-GPU server exceeds it
Three limits, in the order they bite
Every request for a “server with as many GPUs as possible” runs into the same three walls, and they arrive in a fixed order: the PCIe lanes the processors provide, the power the chassis can deliver and dissipate, and the air the rack can move. Vendors design around all three; the customer usually only sees the first. This article walks through what the specification sheets actually say.
Limit one: PCIe lanes
A data-centre GPU is a PCIe device that wants a full x16 link. On Gen5 that link carries about 64 GB/s in each direction; Lenovo quotes 63 GB/s for its x16 slots. NVIDIA’s certified-systems configuration guide (the 2023 PDF revision) is explicit: one Gen5 or Gen4 x16 link per GPU is strongly recommended, x8 is supported but performance may suffer, and at most two GPUs should share one upstream x16 link; the current web version simply says Gen5 x16 or better.
| PROCESSOR | PCIe 5.0 LANES PER SOCKET | TWO-SOCKET SYSTEM |
|---|---|---|
| AMD EPYC 9005 (Turin) | 128 | up to 160 |
| AMD EPYC 9004 (Genoa) | 128 | up to 160 |
| Intel Xeon 6900P | 96 | up to 192 (Intel) |
| Intel Xeon 6700P | 88 | 176 |
| Intel Xeon 5th Gen (8592+) | 80 | 160 |
Sources: AMD EPYC 9005 datasheet and 9004 tuning guide, Intel ARK and Xeon 6 product brief.
The AMD numbers hide a detail worth knowing. In a two-socket EPYC system, each processor lends 64 of its 128 lanes to the socket-to-socket Infinity Fabric links, leaving 64 for devices; drop the fabric from four links to three and each socket keeps 80, hence “up to 160”. Every lane also has to serve NVMe drives, the network cards and the management controller, so the usable count for GPUs is always below the headline.
Once the lanes run out, the vendor adds a PCIe switch. The Dell PowerEdge XE7745 uses four Broadcom Atlas II switches to feed eight double-width slots from two CPUs; the Supermicro 4U line runs a dual-root switch fabric behind thirteen x16 slots; one 2U design fits sixteen single-slot cards behind two switches. NVIDIA allows one or two switch layers and wants network cards and NVMe on the same root complex as the GPUs they feed. The switch is not a compromise for inference: a card reads its own memory, not the bus. It matters when several cards without NVLink exchange tensors, because then two of them may be sharing one uplink.
What that means per chassis height
| CHASSIS | GPUs | SWITCH | POWER SUPPLIES | NOTE |
|---|---|---|---|---|
| 1U Supermicro 121H-TNR | 1 DW or 3 SW | no | 2 × 1,200 W | compute node, one accelerator |
| 2U Supermicro 221GE-NR | 4 DW | no | 3 × 2,000 W | x16 per GPU from the CPUs; L40S, H100 NVL, RTX PRO 6000 listed |
| 2U Dell PowerEdge R760xa | 4 DW up to 400 W, or 12 SW | no | 2,400–3,200 W | 600 W cards exceed the 400 W envelope |
| 2U Lenovo SR650a V4 | 4 DW | no | up to 2 × 3,200 W | 4 × L40S at 350 W, but 2 × H200 NVL at 600 W |
| 3U Lenovo SR675 V3 | 8 DW at 600 W | yes | 4 × 2,600 W, 220 V | RTX PRO 6000 × 8 with 5th-gen EPYC only |
| 4U Dell PowerEdge XE7745 | 8 DW at 600 W, or 16 SW | yes, 4 × Atlas II | 8 bays, 3,200 W Titanium | 2,400 W supplies not supported for 8 × 600 W |
| 4U HPE DL380a Gen12 | up to 10 DW; RTX PRO 6000 max 8 | yes | 8 × 2,400/3,200 W for 8 or 10 | fan upgrade recommended for 600 W cards |
| 4U Supermicro 421GE-TNRT | up to 10 DW | yes, dual-root | 4 × 2,700 W | 13 PCIe 5.0 x16 slots |
| 4U Supermicro AS-4125GS-TNRT | 8 DW | no | 4 × 2,000 W (3+1) | two EPYC sockets, GPUs on the CPUs directly; check the card list for 600 W parts |
DW = double-width card, SW = single-width. Sources: vendor specification sheets and product guides, September 2026. Configurations change with firmware and CPU generation; the vendor’s current guide wins.
The table shows the pattern: up to four double-width cards ride directly on the CPUs in 2U; six to ten need 3U or 4U and switches; and the same 2U accepts four 350 W cards but only two 600 W ones (Lenovo allows four RTX PRO 6000 in the SR650a V4 only when they are capped to 450 W). The card’s power rating, not its length, decides the count more often than any other single line.
Limit two: power
Eight cards at 600 W draw 4.8 kW before the two processors, at 350 to 500 W each, the fans, the memory and the network cards. Dell’s own budget for the XE7745 lists 2,160 W for the GPU-zone fans alone. A fully loaded eight-GPU PCIe server lands in the region of six to eight kilowatts at the wall, which is why every vendor in the table answers with four to eight supplies of 2,000 to 3,200 W, which deliver their full rating only on 200 to 240 V. Dell’s 3,200 W unit derates to 2,900 W at 200 to 220 V, and where a low-line mode exists at all (Dell’s 2,400 W unit falls to 1,400 W on 100 to 120 V) no GPU configuration can use it.
Now the rack. The supplies take C19 or C20 inlets rated at 16 A; a typical 16 A single-phase PDU delivers 3.7 kW and has two C19 outlets. One eight-GPU server exceeds a 16 A feed on its own. This is the line item that most often stops a delivery from being switched on: the server arrives, the rack has 16 A, and the electrician is booked for next month. Plan for 32 A or three-phase feeds, and plan for it before the order.
The card itself needs the right cable. Current 600 W cards use the 12V-2×6 connector from the PCIe CEM 5.1 specification, with sense pins that tell the card whether it may draw 150, 300, 450 or 600 W. Dell documents a specific cable per riser for L40S and H100-class cards and limits some 2U configurations to two double-width cards with power cables. A 600 W card on a 450 W cable runs at 450 W and reports it; it does not fail, it just quietly loses a quarter of its power budget.
Limit three: air
The rule of thumb for air is about 155 CFM per kilowatt at an 11 °C temperature rise across the server, or roughly 90 CFM per kilowatt if the exhaust may run 20 °C above the inlet. A 7 kW server therefore moves in the region of 600 to 1,100 CFM, which is why the GPU-zone fans have their own 2 kW budget.
Inlet temperature is where the fine print lives. ASHRAE class A2, which Dell specifies for the XE7745, allows 10 to 35 °C. Supermicro and the other 4U vendors quote 10 to 35 °C. Lenovo tells SR675 V3 owners to keep ambient at 30 °C or below with L40S or H100 PCIe cards, and at 25 °C with one high-frequency EPYC option. The card adds its own condition: a passive 600 W card needs roughly 70 per cent more airflow than a 350 W L40S at the same inlet, and every degree of inlet temperature costs fan power. A warm hall does not stop the server from booting; it makes the fans work harder, the power budget larger, and the card’s clocks lower.
Where NVLink changes the answer
Among the PCIe cards in this class, the Hopper cards bridge to their neighbours: the H100 PCIe and H100 NVL two-way at 600 GB/s, the H200 NVL two- or four-way NVLink at 900 GB/s per GPU. The RTX PRO 6000 Server Edition and the L40S have no NVLink; every byte between cards crosses PCIe at about 64 GB/s per direction. For inference of one model on one card that is irrelevant. For a model split across four cards it is the whole story, and it is the reason the four-card H200 NVL configuration exists as a product while a four-card RTX PRO 6000 is simply four servers in one box.
How we actually size it
Our engineering partner Vixen.UNO starts from the model, the precision and the concurrency, which give the GPU count. Then the chassis: direct-attached for up to four cards, switched for more. Then the power at the rack, in amps and phases, not kilowatts. Then the inlet temperature the hall can guarantee. The order is deliberate: the last two are the ones that cannot be fixed with a purchase order, and the ones we ask about first.
FAQ
How many GPUs can a 2U server hold?
Do I need a PCIe switch?
Is x8 enough for a GPU?
What does an eight-GPU server draw?
Can I run it on a 16 A circuit?
How much airflow does it need?
Tell us the cards and the rack you have in mind, and we will check lanes, power and air before anything is ordered. We reply within one business day.
Talk to an expertWe reply within one business day