BLOG · GUIDE · 14 SEPTEMBER 2026

Putting an H200 NVL into a server you already own: the checks to do before you order

IN BRIEF
  • The 16-pin cable has to declare the 451 to 600 W class: NVIDIA states that if the level identified by Sense0 and Sense1 is below the card’s default power cap the card will not boot, and every lower class in the table is marked not supported
  • Dell caps the card at 450 W rather than 600 W on both the R770 and the R7725, and marks it unsupported with 24 × 2.5-inch SAS/SATA and with 16, 32 and 40-bay EDSFF E3.S backplanes: the drive configuration decides whether the GPU goes in at all
  • No vendor publishes an airflow rate for this card. NVIDIA’s chapter on system airflow gives only a direction, so the thermal specification is the OEM’s list: named fan modules, a named air duct, permitted slots and an ambient cap of 30 to 35 °C
  • One NVLink bridge per card, not the three an H100 NVL used: a 2-way bridge joins two adjacent cards and a 4-way joins four, both at 900 GB/s per GPU, and every bridged card must sit under the same CPU
  • Driver floor R565 TRD1 and CUDA 12.7; the five-year NVIDIA AI Enterprise subscription is per GPU and activates by serial number, so the serials come off the boards before they are racked. Resizable BAR, by contrast, is not a documented requirement for this card

Gate 1: the slot, and how many of them you really have

The card is 141 GB of HBM3e at 4.8 TB/s on a full-height, full-length, dual-slot PCIe board, ten and a half inches long, drawing up to 600 W through one 16-pin auxiliary connector. Each of those words is a gate, and a retrofit has to pass seven of them in order. A free slot is the easiest of the seven.

ITEMFIGUREWHAT IT DECIDES
Form factorfull-height, full-length, dual-slot, 10.5 intwo slot positions and a full-length bay, not one free slot
Measured envelope10.5 × 1.37 × 4.37 in (266.7 × 34.8 × 111.0 mm)1.37 in is the two-slot pitch; measure the riser cage against it
With the enhanced extender312 mm longa cage built around a 267 mm card will not close
Board weight1,217 g without bracket, extender or bridgewhat the riser cage and retention bracket carry
Host linkGen5 x16, Gen5 x8 or Gen4 x16NVIDIA accepts all three; Gen5 x16 is 128 GB/s

NVIDIA H200 NVL Product Brief PB-12128-001_v01, April 2025, and HPE QuickSpecs NVIDIA Accelerators for HPE, c04123180 v73, July 2025, which measured the same board. The enhanced extender is an option, not the default fitting.

Then count the slots the way the vendor counts them. Dell allows two H200 NVL in a PowerEdge R770 and names the risers that must be present, RC 6-2 and RC 11-2; Lenovo allows two in a ThinkSystem SR650a V4 and only in slots 21 and 23; HPE offers double-wide GPUs in a DL380a Gen12 in quantities of two, four or eight. The usable count is a property of the riser and cable set, not of the empty space at the back of the chassis.

The host is the last mechanical check. NVIDIA’s reference architecture wants one Gen5 x16 link per GPU, tolerates one per two, and sets a floor for the rest of the machine: two sockets, 2.0 GHz, seven physical cores per GPU and 128 GB of system memory per GPU. A server with 256 GB of RAM sits on that floor with two cards and below it with three.

Gate 2: power, and the pin that decides whether the card boots

This gate carries the most concrete failure mode in any of the documents. The card takes up to 600 W from the 16-pin auxiliary connector; the slot itself contributes little, and Dell puts slot power at up to 75 W without an auxiliary cable. That cable carries sense pins declaring what the supply side will deliver, and NVIDIA’s wording is unambiguous: if the power level identified by Sense0 and Sense1 is less than the default power cap of the card, the card will not boot.

POWER LEVEL THE CABLE DECLARESSTATUS ON H200 NVL
451 to 600 W (Sense0 = 0, Sense1 = 0)supported, and the only supported level
301 to 450 Wnot supported
151 to 300 Wnot supported
Up to 150 Wnot supported
No powernot supported

Product Brief PB-12128-001_v01, sense-pin table, labelled PCIe CEM 5.1. Only the highest strap is supported: there is no lower-power cabling option for this card.

A cable strapped for a lower class still mates perfectly, and the result is not a throttled card. It is a card that does not come up at all, in a server that boots normally. The cable is also chassis-specific and slot-specific: Dell names DCNHT for slot 2 and K8FW0 for slot 7 of the same R770, and Lenovo names a 320 mm cable for the SR650a V4 and a 235 mm one for the SR675 V3. Order it by server and slot, by part number.

The configurable range is 600 W by default, with a 350 W power compliance limit and a 200 W minimum in the same table. The 350 W figure is a compliance limit inside the 600 W row, not a variant you can order. A cap set in band with nvidia-smi is lost on the next driver load; only the out-of-band SMBPBI setting survives a reboot.

Then do the arithmetic in front of the power team. Two cards at Dell’s 450 W cap is 900 W; two at Lenovo’s 600 W is 1,200 W. Four at 600 W is 2,400 W, the entire output of one 2,400 W supply before processors, memory, drives and fans have drawn anything, and eight is 4,800 W. HPE’s DL380a Gen12 QuickSpecs accordingly require 2,400 W or 3,200 W supplies whenever the card is configured, five power supplies for a two- or four-GPU node, two for the system board and three for the GPUs, and eight for an eight- or ten-GPU node. Once GPU draw exceeds one supply’s output, N+N stops meaning two supplies. Dell’s instruction for the R770 is to verify consumption with its planning tool rather than a rule of thumb, and at rack level the feed runs out long before the U-space does, which is the arithmetic in our rack power and cooling guide.

Gate 3: cooling, where no vendor publishes a number

The honest answer to “what airflow does this card need” is that nobody publishes one. The heat sink is passive and the card carries no fan, yet NVIDIA’s product brief has a chapter titled System Airflow Requirements whose entire technical content is that the sink accepts air left to right or right to left, plus a figure. No rate in linear feet per minute, none in cubic feet per minute, there or in the equivalent chapter of the H100 PCIe brief. Lenovo’s product guide, HPE’s QuickSpecs, Dell’s technical guide and thermal matrices and Supermicro’s system-guidance paper state none either.

So “our racks have good airflow” is not an answer: there is no published figure to check it against, and an LFM number from a forum or a reseller page has no primary source behind it. NVIDIA publishes no maximum die temperature or throttle threshold either. What the OEMs publish instead is a list of parts and limits, and that list is the thermal specification.

PLATFORMMANDATORY PARTSAMBIENT CAPPOWER CAP
Dell PowerEdge R770all fan modules HPR Platinum type; GPU air shroud35 °C450 W
Dell PowerEdge R7725dual-width GPU shroud, or the card is listed as not applicable30 or 35 °C by drive configuration450 W
Lenovo ThinkSystem SR650a V42U 6056 24K Ultra Fan Module for 600W PCIe Adapter, 2U Front Double Width Air Duct for 600w, slots 21 and 23 only30 °C600 W
The card on its ownpassive sink, no fan, air left to right or right to left45 °C600 W default

Dell R770 and R7725 thermal restriction matrices and the R770 Technical Guide; Lenovo SR650a V4 thermal rules; Product Brief PB-12128-001_v01 Table 2-4. Not one states an airflow rate.

Read the last row against the others. The card runs in ambient air up to 45 °C; the servers certified to hold it do not, and cap the room at 30 or 35 °C. The server number is the one that binds, and note which way the restriction runs on the Dell platforms: the cooling limits the card, not the other way round.

Gate 4: is the server on anyone’s supported list

Four primary documents settle this, and they are worth naming: Lenovo Press LP1944, the ThinkSystem NVIDIA H200 141GB GPUs Product Guide; the HPE QuickSpecs NVIDIA Accelerators for HPE, c04123180 v73 of July 2025; the Dell PowerEdge R770 Technical Guide, part number E111S, Rev. A08 of May 2026; and the Dell PowerEdge Server GPU Matrix. Behind them sit the HPE ProLiant Compute DL380a Gen12 QuickSpecs and Supermicro’s System Guidance for GH200 and H200.

The lists are shorter than most people expect. Dell shows the card on seventeenth-generation systems only, at a maximum of two in the R770 and the R7725 and three in the R7715; it is absent from the fifteenth and sixteenth-generation tables, so an R750 or R760 already in the rack is not on Dell’s list. HPE lists the ProLiant DL380a Gen12 and the ProLiant Compute DL385 Gen11; Lenovo lists the SR650a V4 at two and the SR675 V3 at eight. NVIDIA frames the category as MGX H200 NVL partner and NVIDIA-Certified Systems with up to eight GPUs.

Lenovo’s guide carries the trap that shows how fine-grained this is: supported on the SR650a V4, machine type 7DGDCTO2WW, and explicitly not supported on the SR650 V4, machine type 7DGDCTO1WW. Order against the family name rather than the machine type and it is a coin toss.

On a field installation against a factory build, the documents are less dramatic than the folklore. None says the PCIe card is factory-install only, and none states a warranty consequence for fitting one in the field, so we will not claim either. Lenovo marks the HGX H200 baseboards configure-to-order only, while the PCIe adapter is a normal orderable option on ServerProven. What they do say is that support is scoped to a configuration rather than to a server: a machine type, a slot, a riser, a fan module, an air duct, a cable per slot and a drive backplane on the supported list. HPE puts the difference in qualification terms, saying accelerators sold for NVIDIA-Certified HPE servers undergo thermal, mechanical, power and signal-integrity qualification.

One vendor makes the point sharply. Dell caps the H200 NVL at 450 W rather than its nominal 600 W on two of its own platforms, the R770 and the R7725, and its R770 thermal restriction matrix marks the card not supported with whole classes of drive backplane: 24 × 2.5-inch SAS/SATA, 16, 32 and 40-bay EDSFF E3.S NVMe, and rear E3.S. The backplane sets the air impedance in front of the card, so the disks chosen two years ago can veto the GPU today. Dell also sets an operating-system floor on the R770, Ubuntu 24.04.02 with kernel 6.11 or later, and its generic rules still apply: all GPUs the same type and model, high-performance fans and GPU air shroud fitted.

Gate 5: the NVLink bridge, if the workload needs one

Most retrofits do not need a bridge: if each card holds its own copy of the model, the cards never talk to each other. The bridge earns its place when a model is split across GPUs at every layer, because tensor parallelism inserts an all-reduce per layer and PCIe Gen5 at 128 GB/s is then the bottleneck.

The topology rules are strict. Each H200 NVL carries exactly one bridge connector; the 2-way bridge joins two cards, the 4-way joins four, and in both cases the cards must be physically adjacent and in the same CPU domain. NVIDIA’s reference architecture repeats that pairing under one socket is best and across sockets acceptable but not recommended. In a server whose risers split four GPUs two and two across the sockets, the 4-way bridge is not installable as intended. The chassis must also leave at least 2.5 mm of clearance above the north edge of the card and 2.67 mm behind it.

H200 NVL, 2-WAYH200 NVL, 4-WAYH100 NVL
Bridges per card113
Cards joined2, adjacent4, adjacent2, adjacent
Per-GPU NVLink bandwidth900 GB/s900 GB/s600 GB/s
Aggregate across the set900 GB/s1,800 GB/snot stated in these terms
Memory in the set282 GB564 GB188 GB
Bridge weight49 g128 ga different bridge, not interchangeable

Product Brief PB-12128-001_v01 Tables 4-1 and 4-2, and the H100 NVL Product Brief PB-11773-001_v01. NVIDIA labels the parts “2-slot” and “4-slot” while describing them as spanning two and four cards; since the card is itself dual-slot, read the labels as card counts and derive no pitch from them.

Keep the two bandwidth figures apart, because NVIDIA’s own material blurs them. Per-GPU NVLink bandwidth is 900 GB/s and it is the same on both bridges: 18 links at 50 GB/s per lane in each direction. The 1,800 GB/s printed against the 4-way part is the aggregate across four cards, and 4 × 141 GB gives the 564 GB of pooled memory NVIDIA quotes for a four-way set. NVIDIA’s web datasheet writes the row as 900 GB/s per GPU while the PDF version drops the words “per GPU”, which is where the confusion starts.

This gate changes most for anyone upgrading from an H100 NVL, and one figure needs correcting on the way. NVIDIA’s corporate blog calls the H200 NVL a 1.5 times memory increase and a 1.2 times bandwidth increase over H100 NVL; its developer blog, three weeks later, says 1.4 times bandwidth. The datasheets settle it: 4.8 TB/s divided by 3.9 TB/s is about 1.22, so the uplift is 1.2 times and the 1.4 figure is an error. The interconnect uplift is the larger one, 900 against 600 GB/s, or 1.5 times. The old card also used three bridges per pair where this one uses a single bridge, and HPE states that H200 NVL bridges work only with H200 NVL GPUs, so an existing set does not carry over. A bridged H100 NVL pair is 2 × 94 GB, or 188 GB, which is where the old “188 GB H100 NVL” figure comes from; the equivalent H200 NVL pair is 282 GB, as our comparison of the two cards works through. One last caution: “up to eight GPUs” is a platform statement, not a bridge statement, because the NVLink domain stops at four cards.

Gates 6 and 7: firmware, BIOS and the software that follows

Two firmware floors are documented and both are easy to miss on a machine that has run for two years. For Secure Boot, NVIDIA names CEC firmware 2.0185 or later and NVFlash 5.842 or later; for the platform, NVIDIA-Certified Systems 2.8 or later.

The BIOS setting that decides whether the card initialises is Above 4G decoding, and NVIDIA’s knowledge base warns that vendors name it differently: 64-bit MMIO, Memory Hole for PCI MMIO, Above 4G Decoding. When it is wrong the symptoms are a failure to initialise NVML with an unknown error, or a message that the PCI I/O region assigned to the device is invalid. On some Supermicro boards the option is enabled globally but excluded per slot, so a setting marked enabled can be off for exactly the slot the GPU is in. On Dell sixteenth-generation and later systems the MMIO base knob is gone and the default is 2,048 TB; on Lenovo the knob is MM Config Base, and the documented fix for insufficient PCI resources is counter-intuitively to lower it.

Now the correction, because this is where retrofit checklists go wrong. Resizable BAR is widely repeated as an H200 NVL requirement and it is not one. No NVIDIA or OEM document names it as a prerequisite for this card, and none publishes the card’s BAR1 size, so there is nothing to verify and nothing to enable. The setting engineers are reaching for is the server-side one above: 64-bit MMIO decoding, and enough MMIO space.

If the card will be virtualised, two more settings become requirements. NVIDIA’s vGPU documentation states that VT-d or IOMMU must be enabled and SR-IOV switched on in advanced options; the card needs SBIOS cooperation for SR-IOV and exposes 32 virtual functions. For pass-through, 141 GB of frame buffer breaks MMIO defaults that worked for 24 and 48 GB cards. Broadcom’s rule is to set pciPassthru.use64bitMMIO true, total the frame buffer of every GPU attached to the VM, round up to the next power of two and round up again. One card at 141 GB lies between 128 and 256, so the rule gives 256 and then 512; two cards are 282 GB, which gives 512 and then 1,024.

One conflict cannot be configured away. NVIDIA’s NCCL troubleshooting guide says IO virtualisation can redirect peer-to-peer PCI traffic through the CPU root complex, causing significant performance loss or a hang, and that the bare-metal fix is to disable it; check with lspci -vvv | grep ACSCtl for SrcValid+. The same page states that virtual machines require ACS to function, so a host gives you VM isolation or unrestricted bare-metal peer-to-peer, not both.

The software floors are short: driver R565 TRD1 or later, CUDA 12.7 or later, and vGPU 18.1 or later in the Virtual Compute Server edition. NVIDIA does not designate a current branch for this card in its driver release notes, so treat the branch as a lifecycle decision: of those listed in the September 2026 data centre driver document, R580 LTS has the longest published life, to June 2028, against March 2027 for R595 Production and August 2026 for R610.

Finally the entitlement, which is what goes wrong after everything else has gone right. Every H200 NVL carries a five-year NVIDIA AI Enterprise subscription, the licence asymmetry against the SXM version where the same software is an add-on. It is licensed per GPU rather than per server, and activated by registering the GPU serial numbers on NGC. Nothing happens until somebody does that, and the serials are printed on boards about to disappear into a chassis in a rack. Read them off before the cards go in.

The checks that fail most often

Three of the gates above account for most of the failures, and they fail in three different places.

  1. The 16-pin cable is strapped for the wrong power class. The card does not boot, in a server that does. It fails at first power-on, it looks like a dead GPU, and it is a wrong part number rather than a fault.
  2. The chassis is not on the OEM’s thermal list for this card. This fails at order and validation time, not at boot, and no amount of fan tuning fixes it, because the list is the specification. The drive backplane is the item people forget to check.
  3. MMIO and Above 4G decoding, especially once a hypervisor is involved. A 141 GB card makes this far likelier than the 24 to 48 GB cards most estates were sized for, and the per-slot exclusion is the version of the bug that survives a check of the global setting.

Eurokommerz supplies the NVIDIA H200 NVL and the servers certified to hold it across the EU, and with our engineering partner Vixen.UNO we run this checklist against your machine before anything is ordered: machine type against the vendor list, backplane against the thermal matrix, riser and cable part numbers per slot, power supply count, bridge topology, and the settings that have to change. Where the answer is that the server cannot take the card, we would rather say so before the purchase order than after.

FAQ

Can I fit an H200 NVL into a server I already own?
Only if the exact machine type is on the OEM’s list with the drive configuration you already have. Dell lists the card on seventeenth-generation PowerEdge only, so an R750 or R760 is not on it; Lenovo supports it on the SR650a V4, machine type 7DGDCTO2WW, and explicitly not on the SR650 V4, machine type 7DGDCTO1WW. No vendor document says the card must be factory-fitted, but every one of them scopes support to a configuration rather than to a server.
How much power does an H200 NVL draw, and what power supplies does it need?
Up to 600 W through the 16-pin auxiliary connector, configurable down to a 200 W floor, but the server usually sets the real number: Dell caps it at 450 W on both the R770 and the R7725. HPE requires 2,400 W or 3,200 W supplies whenever the card is configured, five of them for a two- or four-GPU node and eight for an eight- or ten-GPU node. Size with the vendor’s planning tool rather than with a rule of thumb.
What airflow does an H200 NVL require in LFM?
No vendor publishes one. NVIDIA’s product brief has a chapter called System Airflow Requirements whose only technical content is that the passive sink accepts air left to right or right to left, and the Dell, HPE, Lenovo and Supermicro documents state no rate either. Treat any linear-feet-per-minute figure you find for this card as unsourced unless it comes from your chassis vendor’s own thermal data.
Do I need an NVLink bridge for two H200 NVL cards?
Only if a model is split across both GPUs at every layer. Each card carries one bridge connector, the 2-way bridge joins two physically adjacent cards at 900 GB/s per GPU, and both cards must sit under the same CPU. Data-parallel serving, where each card holds its own copy of the model, needs no bridge at all, and an H100 NVL three-bridge set does not carry over.
Does the H200 NVL need Resizable BAR enabled?
No NVIDIA or OEM document names Resizable BAR as a requirement for this card, and none publishes its BAR1 size. The setting that matters is Above 4G decoding, also called 64-bit MMIO or Memory Hole for PCI MMIO depending on the vendor, together with enough MMIO space. On some Supermicro boards, check the per-slot exclusion as well, because Above 4G can be on globally and off for the slot the card is in.
Is NVIDIA AI Enterprise included with the H200 NVL?
Yes, a five-year subscription, which is the licence difference from the SXM version where the same software is an add-on. It is licensed per GPU rather than per server and it is activated by registering the GPU serial numbers on NGC. Read the serials off the boards before they are racked; an entitlement nobody activated is the most common post-installation surprise.

Send us the server model, the machine type, the drive configuration and the slot layout, and we will tell you whether the card fits before anyone raises an order. We reply within one business day.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at  ·  +43 1 585 1405 50  ·  Jordangasse 7, 1010 Vienna