Putting an H200 NVL into a server you already own: the checks to do before you order
- The 16-pin cable has to declare the 451 to 600 W class: NVIDIA states that if the level identified by Sense0 and Sense1 is below the card’s default power cap the card will not boot, and every lower class in the table is marked not supported
- Dell caps the card at 450 W rather than 600 W on both the R770 and the R7725, and marks it unsupported with 24 × 2.5-inch SAS/SATA and with 16, 32 and 40-bay EDSFF E3.S backplanes: the drive configuration decides whether the GPU goes in at all
- No vendor publishes an airflow rate for this card. NVIDIA’s chapter on system airflow gives only a direction, so the thermal specification is the OEM’s list: named fan modules, a named air duct, permitted slots and an ambient cap of 30 to 35 °C
- One NVLink bridge per card, not the three an H100 NVL used: a 2-way bridge joins two adjacent cards and a 4-way joins four, both at 900 GB/s per GPU, and every bridged card must sit under the same CPU
- Driver floor R565 TRD1 and CUDA 12.7; the five-year NVIDIA AI Enterprise subscription is per GPU and activates by serial number, so the serials come off the boards before they are racked. Resizable BAR, by contrast, is not a documented requirement for this card
Gate 1: the slot, and how many of them you really have
The card is 141 GB of HBM3e at 4.8 TB/s on a full-height, full-length, dual-slot PCIe board, ten and a half inches long, drawing up to 600 W through one 16-pin auxiliary connector. Each of those words is a gate, and a retrofit has to pass seven of them in order. A free slot is the easiest of the seven.
| ITEM | FIGURE | WHAT IT DECIDES |
|---|---|---|
| Form factor | full-height, full-length, dual-slot, 10.5 in | two slot positions and a full-length bay, not one free slot |
| Measured envelope | 10.5 × 1.37 × 4.37 in (266.7 × 34.8 × 111.0 mm) | 1.37 in is the two-slot pitch; measure the riser cage against it |
| With the enhanced extender | 312 mm long | a cage built around a 267 mm card will not close |
| Board weight | 1,217 g without bracket, extender or bridge | what the riser cage and retention bracket carry |
| Host link | Gen5 x16, Gen5 x8 or Gen4 x16 | NVIDIA accepts all three; Gen5 x16 is 128 GB/s |
NVIDIA H200 NVL Product Brief PB-12128-001_v01, April 2025, and HPE QuickSpecs NVIDIA Accelerators for HPE, c04123180 v73, July 2025, which measured the same board. The enhanced extender is an option, not the default fitting.
Then count the slots the way the vendor counts them. Dell allows two H200 NVL in a PowerEdge R770 and names the risers that must be present, RC 6-2 and RC 11-2; Lenovo allows two in a ThinkSystem SR650a V4 and only in slots 21 and 23; HPE offers double-wide GPUs in a DL380a Gen12 in quantities of two, four or eight. The usable count is a property of the riser and cable set, not of the empty space at the back of the chassis.
The host is the last mechanical check. NVIDIA’s reference architecture wants one Gen5 x16 link per GPU, tolerates one per two, and sets a floor for the rest of the machine: two sockets, 2.0 GHz, seven physical cores per GPU and 128 GB of system memory per GPU. A server with 256 GB of RAM sits on that floor with two cards and below it with three.
Gate 2: power, and the pin that decides whether the card boots
This gate carries the most concrete failure mode in any of the documents. The card takes up to 600 W from the 16-pin auxiliary connector; the slot itself contributes little, and Dell puts slot power at up to 75 W without an auxiliary cable. That cable carries sense pins declaring what the supply side will deliver, and NVIDIA’s wording is unambiguous: if the power level identified by Sense0 and Sense1 is less than the default power cap of the card, the card will not boot.
| POWER LEVEL THE CABLE DECLARES | STATUS ON H200 NVL |
|---|---|
| 451 to 600 W (Sense0 = 0, Sense1 = 0) | supported, and the only supported level |
| 301 to 450 W | not supported |
| 151 to 300 W | not supported |
| Up to 150 W | not supported |
| No power | not supported |
Product Brief PB-12128-001_v01, sense-pin table, labelled PCIe CEM 5.1. Only the highest strap is supported: there is no lower-power cabling option for this card.
A cable strapped for a lower class still mates perfectly, and the result is not a throttled card. It is a card that does not come up at all, in a server that boots normally. The cable is also chassis-specific and slot-specific: Dell names DCNHT for slot 2 and K8FW0 for slot 7 of the same R770, and Lenovo names a 320 mm cable for the SR650a V4 and a 235 mm one for the SR675 V3. Order it by server and slot, by part number.
The configurable range is 600 W by default, with a 350 W power compliance limit and a 200 W minimum in the same table. The 350 W figure is a compliance limit inside the 600 W row, not a variant you can order. A cap set in band with nvidia-smi is lost on the next driver load; only the out-of-band SMBPBI setting survives a reboot.
Then do the arithmetic in front of the power team. Two cards at Dell’s 450 W cap is 900 W; two at Lenovo’s 600 W is 1,200 W. Four at 600 W is 2,400 W, the entire output of one 2,400 W supply before processors, memory, drives and fans have drawn anything, and eight is 4,800 W. HPE’s DL380a Gen12 QuickSpecs accordingly require 2,400 W or 3,200 W supplies whenever the card is configured, five power supplies for a two- or four-GPU node, two for the system board and three for the GPUs, and eight for an eight- or ten-GPU node. Once GPU draw exceeds one supply’s output, N+N stops meaning two supplies. Dell’s instruction for the R770 is to verify consumption with its planning tool rather than a rule of thumb, and at rack level the feed runs out long before the U-space does, which is the arithmetic in our rack power and cooling guide.
Gate 3: cooling, where no vendor publishes a number
The honest answer to “what airflow does this card need” is that nobody publishes one. The heat sink is passive and the card carries no fan, yet NVIDIA’s product brief has a chapter titled System Airflow Requirements whose entire technical content is that the sink accepts air left to right or right to left, plus a figure. No rate in linear feet per minute, none in cubic feet per minute, there or in the equivalent chapter of the H100 PCIe brief. Lenovo’s product guide, HPE’s QuickSpecs, Dell’s technical guide and thermal matrices and Supermicro’s system-guidance paper state none either.
So “our racks have good airflow” is not an answer: there is no published figure to check it against, and an LFM number from a forum or a reseller page has no primary source behind it. NVIDIA publishes no maximum die temperature or throttle threshold either. What the OEMs publish instead is a list of parts and limits, and that list is the thermal specification.
| PLATFORM | MANDATORY PARTS | AMBIENT CAP | POWER CAP |
|---|---|---|---|
| Dell PowerEdge R770 | all fan modules HPR Platinum type; GPU air shroud | 35 °C | 450 W |
| Dell PowerEdge R7725 | dual-width GPU shroud, or the card is listed as not applicable | 30 or 35 °C by drive configuration | 450 W |
| Lenovo ThinkSystem SR650a V4 | 2U 6056 24K Ultra Fan Module for 600W PCIe Adapter, 2U Front Double Width Air Duct for 600w, slots 21 and 23 only | 30 °C | 600 W |
| The card on its own | passive sink, no fan, air left to right or right to left | 45 °C | 600 W default |
Dell R770 and R7725 thermal restriction matrices and the R770 Technical Guide; Lenovo SR650a V4 thermal rules; Product Brief PB-12128-001_v01 Table 2-4. Not one states an airflow rate.
Read the last row against the others. The card runs in ambient air up to 45 °C; the servers certified to hold it do not, and cap the room at 30 or 35 °C. The server number is the one that binds, and note which way the restriction runs on the Dell platforms: the cooling limits the card, not the other way round.
Gate 4: is the server on anyone’s supported list
Four primary documents settle this, and they are worth naming: Lenovo Press LP1944, the ThinkSystem NVIDIA H200 141GB GPUs Product Guide; the HPE QuickSpecs NVIDIA Accelerators for HPE, c04123180 v73 of July 2025; the Dell PowerEdge R770 Technical Guide, part number E111S, Rev. A08 of May 2026; and the Dell PowerEdge Server GPU Matrix. Behind them sit the HPE ProLiant Compute DL380a Gen12 QuickSpecs and Supermicro’s System Guidance for GH200 and H200.
The lists are shorter than most people expect. Dell shows the card on seventeenth-generation systems only, at a maximum of two in the R770 and the R7725 and three in the R7715; it is absent from the fifteenth and sixteenth-generation tables, so an R750 or R760 already in the rack is not on Dell’s list. HPE lists the ProLiant DL380a Gen12 and the ProLiant Compute DL385 Gen11; Lenovo lists the SR650a V4 at two and the SR675 V3 at eight. NVIDIA frames the category as MGX H200 NVL partner and NVIDIA-Certified Systems with up to eight GPUs.
Lenovo’s guide carries the trap that shows how fine-grained this is: supported on the SR650a V4, machine type 7DGDCTO2WW, and explicitly not supported on the SR650 V4, machine type 7DGDCTO1WW. Order against the family name rather than the machine type and it is a coin toss.
On a field installation against a factory build, the documents are less dramatic than the folklore. None says the PCIe card is factory-install only, and none states a warranty consequence for fitting one in the field, so we will not claim either. Lenovo marks the HGX H200 baseboards configure-to-order only, while the PCIe adapter is a normal orderable option on ServerProven. What they do say is that support is scoped to a configuration rather than to a server: a machine type, a slot, a riser, a fan module, an air duct, a cable per slot and a drive backplane on the supported list. HPE puts the difference in qualification terms, saying accelerators sold for NVIDIA-Certified HPE servers undergo thermal, mechanical, power and signal-integrity qualification.
One vendor makes the point sharply. Dell caps the H200 NVL at 450 W rather than its nominal 600 W on two of its own platforms, the R770 and the R7725, and its R770 thermal restriction matrix marks the card not supported with whole classes of drive backplane: 24 × 2.5-inch SAS/SATA, 16, 32 and 40-bay EDSFF E3.S NVMe, and rear E3.S. The backplane sets the air impedance in front of the card, so the disks chosen two years ago can veto the GPU today. Dell also sets an operating-system floor on the R770, Ubuntu 24.04.02 with kernel 6.11 or later, and its generic rules still apply: all GPUs the same type and model, high-performance fans and GPU air shroud fitted.
Gate 5: the NVLink bridge, if the workload needs one
Most retrofits do not need a bridge: if each card holds its own copy of the model, the cards never talk to each other. The bridge earns its place when a model is split across GPUs at every layer, because tensor parallelism inserts an all-reduce per layer and PCIe Gen5 at 128 GB/s is then the bottleneck.
The topology rules are strict. Each H200 NVL carries exactly one bridge connector; the 2-way bridge joins two cards, the 4-way joins four, and in both cases the cards must be physically adjacent and in the same CPU domain. NVIDIA’s reference architecture repeats that pairing under one socket is best and across sockets acceptable but not recommended. In a server whose risers split four GPUs two and two across the sockets, the 4-way bridge is not installable as intended. The chassis must also leave at least 2.5 mm of clearance above the north edge of the card and 2.67 mm behind it.
| H200 NVL, 2-WAY | H200 NVL, 4-WAY | H100 NVL | |
|---|---|---|---|
| Bridges per card | 1 | 1 | 3 |
| Cards joined | 2, adjacent | 4, adjacent | 2, adjacent |
| Per-GPU NVLink bandwidth | 900 GB/s | 900 GB/s | 600 GB/s |
| Aggregate across the set | 900 GB/s | 1,800 GB/s | not stated in these terms |
| Memory in the set | 282 GB | 564 GB | 188 GB |
| Bridge weight | 49 g | 128 g | a different bridge, not interchangeable |
Product Brief PB-12128-001_v01 Tables 4-1 and 4-2, and the H100 NVL Product Brief PB-11773-001_v01. NVIDIA labels the parts “2-slot” and “4-slot” while describing them as spanning two and four cards; since the card is itself dual-slot, read the labels as card counts and derive no pitch from them.
Keep the two bandwidth figures apart, because NVIDIA’s own material blurs them. Per-GPU NVLink bandwidth is 900 GB/s and it is the same on both bridges: 18 links at 50 GB/s per lane in each direction. The 1,800 GB/s printed against the 4-way part is the aggregate across four cards, and 4 × 141 GB gives the 564 GB of pooled memory NVIDIA quotes for a four-way set. NVIDIA’s web datasheet writes the row as 900 GB/s per GPU while the PDF version drops the words “per GPU”, which is where the confusion starts.
This gate changes most for anyone upgrading from an H100 NVL, and one figure needs correcting on the way. NVIDIA’s corporate blog calls the H200 NVL a 1.5 times memory increase and a 1.2 times bandwidth increase over H100 NVL; its developer blog, three weeks later, says 1.4 times bandwidth. The datasheets settle it: 4.8 TB/s divided by 3.9 TB/s is about 1.22, so the uplift is 1.2 times and the 1.4 figure is an error. The interconnect uplift is the larger one, 900 against 600 GB/s, or 1.5 times. The old card also used three bridges per pair where this one uses a single bridge, and HPE states that H200 NVL bridges work only with H200 NVL GPUs, so an existing set does not carry over. A bridged H100 NVL pair is 2 × 94 GB, or 188 GB, which is where the old “188 GB H100 NVL” figure comes from; the equivalent H200 NVL pair is 282 GB, as our comparison of the two cards works through. One last caution: “up to eight GPUs” is a platform statement, not a bridge statement, because the NVLink domain stops at four cards.
Gates 6 and 7: firmware, BIOS and the software that follows
Two firmware floors are documented and both are easy to miss on a machine that has run for two years. For Secure Boot, NVIDIA names CEC firmware 2.0185 or later and NVFlash 5.842 or later; for the platform, NVIDIA-Certified Systems 2.8 or later.
The BIOS setting that decides whether the card initialises is Above 4G decoding, and NVIDIA’s knowledge base warns that vendors name it differently: 64-bit MMIO, Memory Hole for PCI MMIO, Above 4G Decoding. When it is wrong the symptoms are a failure to initialise NVML with an unknown error, or a message that the PCI I/O region assigned to the device is invalid. On some Supermicro boards the option is enabled globally but excluded per slot, so a setting marked enabled can be off for exactly the slot the GPU is in. On Dell sixteenth-generation and later systems the MMIO base knob is gone and the default is 2,048 TB; on Lenovo the knob is MM Config Base, and the documented fix for insufficient PCI resources is counter-intuitively to lower it.
Now the correction, because this is where retrofit checklists go wrong. Resizable BAR is widely repeated as an H200 NVL requirement and it is not one. No NVIDIA or OEM document names it as a prerequisite for this card, and none publishes the card’s BAR1 size, so there is nothing to verify and nothing to enable. The setting engineers are reaching for is the server-side one above: 64-bit MMIO decoding, and enough MMIO space.
If the card will be virtualised, two more settings become requirements. NVIDIA’s vGPU documentation states that VT-d or IOMMU must be enabled and SR-IOV switched on in advanced options; the card needs SBIOS cooperation for SR-IOV and exposes 32 virtual functions. For pass-through, 141 GB of frame buffer breaks MMIO defaults that worked for 24 and 48 GB cards. Broadcom’s rule is to set pciPassthru.use64bitMMIO true, total the frame buffer of every GPU attached to the VM, round up to the next power of two and round up again. One card at 141 GB lies between 128 and 256, so the rule gives 256 and then 512; two cards are 282 GB, which gives 512 and then 1,024.
One conflict cannot be configured away. NVIDIA’s NCCL troubleshooting guide says IO virtualisation can redirect peer-to-peer PCI traffic through the CPU root complex, causing significant performance loss or a hang, and that the bare-metal fix is to disable it; check with lspci -vvv | grep ACSCtl for SrcValid+. The same page states that virtual machines require ACS to function, so a host gives you VM isolation or unrestricted bare-metal peer-to-peer, not both.
The software floors are short: driver R565 TRD1 or later, CUDA 12.7 or later, and vGPU 18.1 or later in the Virtual Compute Server edition. NVIDIA does not designate a current branch for this card in its driver release notes, so treat the branch as a lifecycle decision: of those listed in the September 2026 data centre driver document, R580 LTS has the longest published life, to June 2028, against March 2027 for R595 Production and August 2026 for R610.
Finally the entitlement, which is what goes wrong after everything else has gone right. Every H200 NVL carries a five-year NVIDIA AI Enterprise subscription, the licence asymmetry against the SXM version where the same software is an add-on. It is licensed per GPU rather than per server, and activated by registering the GPU serial numbers on NGC. Nothing happens until somebody does that, and the serials are printed on boards about to disappear into a chassis in a rack. Read them off before the cards go in.
The checks that fail most often
Three of the gates above account for most of the failures, and they fail in three different places.
- The 16-pin cable is strapped for the wrong power class. The card does not boot, in a server that does. It fails at first power-on, it looks like a dead GPU, and it is a wrong part number rather than a fault.
- The chassis is not on the OEM’s thermal list for this card. This fails at order and validation time, not at boot, and no amount of fan tuning fixes it, because the list is the specification. The drive backplane is the item people forget to check.
- MMIO and Above 4G decoding, especially once a hypervisor is involved. A 141 GB card makes this far likelier than the 24 to 48 GB cards most estates were sized for, and the per-slot exclusion is the version of the bug that survives a check of the global setting.
Eurokommerz supplies the NVIDIA H200 NVL and the servers certified to hold it across the EU, and with our engineering partner Vixen.UNO we run this checklist against your machine before anything is ordered: machine type against the vendor list, backplane against the thermal matrix, riser and cable part numbers per slot, power supply count, bridge topology, and the settings that have to change. Where the answer is that the server cannot take the card, we would rather say so before the purchase order than after.
FAQ
Can I fit an H200 NVL into a server I already own?
How much power does an H200 NVL draw, and what power supplies does it need?
What airflow does an H200 NVL require in LFM?
Do I need an NVLink bridge for two H200 NVL cards?
Does the H200 NVL need Resizable BAR enabled?
Is NVIDIA AI Enterprise included with the H200 NVL?
Send us the server model, the machine type, the drive configuration and the slot layout, and we will tell you whether the card fits before anyone raises an order. We reply within one business day.
Talk to an expertWe reply within one business day