One 8-GPU server or two 4-GPU servers: H200 NVL and RTX PRO 6000 compared
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- NVIDIA lists 2-way and 4-way NVLink bridges for the H200 NVL and “up to four GPUs connected by NVIDIA NVLink”, so by our reading eight cards form at most two NVLink domains of four, in one server or in two; the RTX PRO 6000 Server Edition has no NVLink
- A host failure, reboot or driver update stops all eight cards of one server, but only four with two servers; in our 2,000-staff example with gpt-oss-120b at 32K, two four-card H200 NVL servers keep 220 conversations with one host down
- A model above the 564 GB of four H200 NVL, such as DeepSeek-V3.1 in FP8 at about 689 GB, needs all eight cards; split into two pipeline stages, it crosses PCIe in one server and the network in two, and either host failure stops it in both layouts
- Eight 600 W cards draw 4.8 kW on their own and need a three-phase feed at one rack position; two four-card servers draw 2.4 kW each for the cards and can sit on separate racks and feeds
- NVIDIA AI Enterprise is licensed per GPU, so eight cards that run its software need eight licences in either layout; each H200 NVL includes a five-year subscription, the RTX PRO 6000 Server Edition none; chassis, power supplies, management controllers and network ports are bought twice
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
One 8-GPU server or two 4-GPU servers: the short answer
One 8-GPU server and two 4-GPU servers give the same GPU memory and need the same number of per-GPU licences. With the H200 NVL they also give the same NVLink layout, because NVIDIA’s bridges join two or four cards, and by our reading eight cards form at most two NVLink domains of four in both cases. The decision rests on the remaining differences: how many cards stop when one host fails or reboots, whether a model needs more than four cards, the power and space each rack position offers, and the chassis, processors and ports you buy twice.
Two four-card servers suit models that fit four cards and a service that has to keep running through a host failure or a maintenance window. One eight-card server suits a model that needs all eight cards, fine-tuning across eight cards, or a site with a single rack position that can take the load.
| CRITERION | ONE SERVER, 8 GPUS | TWO SERVERS, 4 GPUS |
|---|---|---|
| H200 NVL NVLink | two 4-way domains, joined over PCIe | one 4-way domain per server |
| Model above four cards | split inside one host | split across two hosts over the network |
| Host failure or reboot | all eight cards stop | four cards stop, four keep serving |
| Rack space, Lenovo SR675 V3 | 3U | 2 × 3U |
| Power for the cards | 4.8 kW at one position | 2.4 kW per server, separate feeds possible |
| NVIDIA AI Enterprise | eight GPU licences | eight GPU licences |
| Chassis, supplies, ports | one set | two sets |
NVIDIA H200 product page and AI Enterprise licensing guide (2 September 2026), Lenovo Press LP1611 (8 October 2026), all read on 10 October 2026; card power at 600 W per card as NVIDIA and Lenovo rate it. Two domains in an eight-card server is our reading of NVIDIA’s four-card bridge limit.
H200 NVL: NVLink domains of four in both layouts
NVIDIA’s specification for the H200 NVL lists a “2- or 4-way NVIDIA NVLink bridge: 900GB/s per GPU” beside “PCIe Gen5: 128GB/s”, and server options “with up to 8 GPUs”. Its product page speaks of “up to four GPUs connected by NVIDIA NVLink”. We found no NVIDIA document that describes how the bridges are arranged in an eight-card server. A four-card limit and eight cards per server mean at most two domains of four, joined over PCIe, or four pairs with 2-way bridges. Our guide to H200 NVL card counts for large models sizes two domains of four. Lenovo’s 3U ThinkSystem SR675 V3, listed for up to eight double-wide GPUs “with NVLink bridges”, routes its eight-card x16 configurations through a Gen5 PCIe switch.
In two four-card servers, each server holds one complete domain, and the link between the domains is the network instead of PCIe. Inside a domain, tensor parallelism splits every layer over NVLink. Between domains, a common layout splits the model into stages with pipeline parallelism, which sends only the activations at the stage boundary. For a model that fits four cards, the domains never need to talk to each other, and the two layouts behave the same until a host goes down.
Failure domain and maintenance windows
A host is the unit that fails and the unit you maintain. A failed mainboard or processor, a firmware update, a reboot after a kernel or driver update and a GPU driver reload all stop every card in that host. With one eight-card server that is the whole service. With two four-card servers it is half the cards, and the other server keeps serving while the first is in its maintenance window. A single failed card affects both layouts alike at first: replicas on the other cards continue, and a model split across four cards loses that copy. Replacing a PCIe card means powering its server off, and that window again stops eight cards or four.
The capacity left after a host failure shows the difference. Our guide to private ChatGPT servers by company size estimates that one RTX PRO 6000 holds about 19 conversations of gpt-oss-120b at a declared 32K context with a 16-bit cache and one H200 NVL about 55, and its 2,000-staff example has a peak of 80 requests in flight.
| LAYOUT | ALL HOSTS UP | ONE HOST DOWN |
|---|---|---|
| 1 × 8 H200 NVL | 440 | 0 |
| 2 × 4 H200 NVL | 440 | 220 |
| 1 × 8 RTX PRO 6000 | 152 | 0 |
| 2 × 4 RTX PRO 6000 | 152 | 76, below the peak of 80 |
Concurrent conversations of gpt-oss-120b at 32K, one copy per card, with the per-card estimates from our company-size guide; our arithmetic.
Two four-card H200 NVL servers keep the example peak with one host down. Two four-card RTX PRO 6000 servers miss it by four conversations, which is why our guide to two-node high availability for LLM inference sizes five cards per node for this case. The same guide covers health checks, load balancing and rolling updates.
We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users. Tell us your models and how long the service may be down in the form below.
Models that need more than four cards
Some models do not fit one domain. By the figures in our H200 NVL guide, DeepSeek-V3.1 or R1 in FP8 takes about 689 GB and Kimi K2 Thinking in INT4 594 GB, both above the 564 GB of four H200 NVL. Such a model needs all eight cards. Our H200 NVL guide sizes it as two pipeline stages of four, one per NVLink domain; layouts such as tensor parallelism over all eight cards send more traffic between the domains.
In one eight-card server, the stage boundary crosses PCIe inside the host and vLLM runs on one machine. In two four-card servers it crosses the network, and vLLM runs as a multi-node deployment. Its documentation sets the tensor-parallel size to the GPUs per node and the pipeline-parallel size to the number of nodes, notes that such deployments “can use Ray as the runtime engine” and asks that “every node provides an identical execution environment, including the model path and Python packages”. It also warns that “Efficient tensor parallelism requires fast internode communication”, so tensor parallelism stays inside each server. NVIDIA’s configuration guide for NVIDIA-Certified Systems asks for a “Minimum 200 Gbps for multi-node inference”, which is 25 GB/s, against 64 GB/s in each direction for one PCIe 5.0 x16 link.
Availability does not improve with the split. The model needs both servers, so the loss of either host stops it, as the loss of the single eight-card host does. A service that must survive a host failure with a model of this size needs a second set of eight cards. For one copy of such a model, the single eight-card server is the simpler layout, since it needs no cluster runtime and no fast network between servers.
Eight RTX PRO 6000: one PCIe server or two
The RTX PRO 6000 Blackwell Server Edition has 96 GB of GDDR7 and a configurable power limit of up to 600 W, and NVIDIA’s specification lists no NVLink for it. Its cards exchange data over PCIe in both layouts, through the server’s PCIe switches inside one host and over the network between two. In the examples of our company-size guide, gpt-oss-120b runs one copy per card and DeepSeek-V4-Flash one copy per pair of cards, and then the choice is the failure domain from the section above. With DeepSeek-V4-Flash, our company-size guide puts two copies on each four-card server, enough to keep the example peak of 80 if one server fails.
Four RTX PRO 6000 hold 384 GB, so a model above that, such as Qwen3.5-397B in FP8 at 406 GB, needs more than four cards. Keep it in one eight-card server, where vLLM’s documentation suggests pipeline parallelism “if the GPUs on the node do not have NVLINK interconnect”. Our article on splitting one model over PCIe or NVLink covers the traffic of each split and the peer-to-peer settings a PCIe server needs.
Chassis, rack space and power per rack position
Lenovo builds the SR675 V3 in one 3U chassis as a 4-DW and an 8-DW GPU model, both rated for double-wide GPUs at 600 W each, with up to four power supplies of 1,800, 2,400 or 2,600 W on 220 V. In that family, two four-card servers take 6U and one eight-card server 3U. A 2U server does not always take four 600 W cards. Lenovo’s 2U SR650a V4 supports “four 400W or two 600W double-wide GPUs”, so it takes two H200 NVL at 600 W; Lenovo’s RTX PRO 6000 guide adds a version capped at 450 W “to allow 4x GPUs to be installed in the SR650a V4”.
Power is planned per rack position. Eight cards at 600 W draw 4.8 kW before processors, memory and fans, more than the about 3.7 kW of a 16 A single-phase feed at 230 V, so one eight-card server needs three-phase feeds at its position. Four cards draw 2.4 kW, and our article on how many GPUs fit in one server plans 32 A or three-phase from four 600 W cards upwards. Two servers can go into two racks on separate feeds, which also spreads the heat. Where only one rack position has three-phase power, one eight-card server may be the only layout that fits; our guide to GPU rack power and cooling covers the room side.
Processors, memory, ports and licences: what doubles
NVIDIA’s configuration guide, updated on 30 September 2026, calls “2x / 4x / 8x GPUs per server” balanced and starts from two CPU sockets, six physical cores per GPU and system memory of twice the total GPU memory. Cores and memory scale with the cards, so the totals are the same in both layouts: 48 cores, and 2,256 GB for eight H200 NVL or 1,536 GB for eight RTX PRO 6000. What doubles is the platform: at the guide’s minimum four sockets instead of two, two chassis, two sets of power supplies, two management controllers, two boot mirrors, a copy of the model store on each server and twice the network ports.
NVIDIA AI Enterprise is licensed per GPU, and its licensing guide requires a licence “for every GPU installed on the server or workstation” that hosts its software, so eight such cards need eight licences in either layout. Each H200 NVL includes a five-year subscription that starts from the board’s ship date to the server maker, plus 90 days; the RTX PRO 6000 Server Edition includes none. Software licensed per server or per processor, such as a hypervisor or an operating system subscription, follows the number of servers and processors.
Two servers also allow a staged start: one four-card server first and a second of the same build later. Growing one server from four cards to eight works only if it was ordered for eight, as our article on planning a GPU server for growth explains.
Which layout fits which workload
| WORKLOAD | LAYOUT | WHY |
|---|---|---|
| Chat model on one card | two × 4 | half the cards keep serving during a failure or update |
| 355B to 405B class in FP8 | two × 4 H200 NVL | one copy per server on its 4-way bridge |
| 671B FP8 or 1T INT4 | one × 8 H200 NVL | the model needs both NVLink domains |
| Above 384 GB on RTX PRO 6000 | one × 8 RTX PRO 6000 | pipeline stages over PCIe, not the network |
| Fine-tuning on eight cards | one × 8 | gradients cross PCIe at every step |
| One three-phase position | one × 8 | 4.8 kW of cards at one position |
Model sizes from our H200 NVL guide (Hugging Face, September 2026); our reading of the sections above. A 355B to 405B model in FP8, such as GLM-4.5 at 361 GB, fits the 564 GB of four H200 NVL by its weights; its cache sets the number of conversations.
We check the rack, power and airflow before we quote. Describe your rack positions, their feeds and the models you plan in the form below.
What we supply
We build AI servers to order with four or eight H200 NVL and their two-way and four-way NVLink bridges, or with the RTX PRO 6000 Server Edition, assembled and burn-in tested, with manufacturer warranty on every component. Both layouts come on one EU contract and invoice, with NVIDIA AI Enterprise and vGPU licences on the same invoice. We also supply the cards on their own for servers you already run, after a compatibility check of the platform, power and cooling. Send us the models and the uptime you need, and within one business day we return a configuration and quote for the layout that fits.
FAQ
Is one 8-GPU server better than two 4-GPU servers?
Can eight H200 NVL be connected with one NVLink bridge?
How much memory do four H200 NVL have?
What happens when one GPU server fails?
How much power does a 4-GPU or an 8-GPU server need?
Do two 4-GPU servers need more NVIDIA AI Enterprise licences than one 8-GPU server?
Send us the models you plan to run with their precision and context length, the peak number of requests in flight, how long the service may be down and the rack positions with their power feeds. We reply within one business day with a configuration and quote for the layout that fits, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day