BLOG · GUIDE ·

GPU server acceptance testing: burn-in, DCGM diagnostics and what to check on delivery

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Accept a GPU server in four steps: identity (every card, its serial, VBIOS and PCIe link), health with DCGM diagnostics, power and temperature under sustained load, and GPU-to-GPU bandwidth where the cards exchange data
  • DCGM diagnostics have four run levels; level 3 adds the diagnostic, targeted stress, targeted power, nvbandwidth and NCCL tests to level 2 and takes under 10 minutes on a 4-GPU Hopper system by the DCGM documentation, while NVIDIA’s triage guide says the long test should take approximately 30 minutes
  • DCGM’s feature table gives every diagnostic level only to Tesla, NVIDIA’s old name for its data-centre line, and level 1 to Quadro, Titan and GeForce; since DCGM 4.6.0, levels 3 and 4 can run as stability checks on GPUs without SKU-specific calibration
  • Read the PCIe link while the GPU is busy, because nvidia-smi says the current generation and width may be reduced when the GPU is not in use; NVIDIA lists Gen5 for the RTX PRO 6000 Server Edition and H200 NVL and Gen4 x16 for the L40S and L4
  • Targeted power drives each GPU towards its TDP and fails below 75 per cent of the target by default; as our acceptance criterion, uncorrectable ECC errors, pending row remaps and Xid messages after the burn-in should all be zero

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What to check when a GPU server is delivered

A GPU server acceptance test, with a burn-in, checks four things before the server goes into production. First confirm its identity, meaning that every ordered card is present with the expected serial number, VBIOS and PCIe link. Then check its health with NVIDIA’s DCGM diagnostics, its power and temperature under sustained load and, on servers where the cards exchange data, the bandwidth between GPUs. The table lists each check, its tool and the pass criterion.

CHECKTOOLPASS CRITERIONBASIS
Cards and serialsnvidia-smi -L, nvidia-smi -qevery ordered card listed; serials match the board labelsNVIDIA
VBIOS and firmwarenvidia-smi -qidentical versions on identical cardsour criterion
PCIe linknvidia-smi -q under loadcurrent generation and width equal the maximumNVIDIA, link drops when idle
Healthdcgmi diag -r 3every test passesNVIDIA DCGM
Sustained powerdcgmi diag -r targeted_powerat least 75 per cent of target powerDCGM default
Clocks and temperaturenvidia-smi -q -d PERFORMANCEno thermal slowdown during the runour criterion
Memory errorsnvidia-smi -q -d ECC,ROW_REMAPPERno uncorrectable errors, no pending remapsDCGM software test
Driver errorskernel logno Xid messages during the runour criterion
GPU-to-GPU bandwidthnvbandwidth, nccl-testscomparable figures on every GPU pairour criterion

nvidia-smi documentation and DCGM 4.6 diagnostics documentation, read on 10 October 2026; rows marked «our criterion» are our examples, not NVIDIA pass thresholds.

We found no NVIDIA document that sets a burn-in duration or calls DCGM an acceptance test. DCGM’s documentation lists among its goals a tool “to assess cluster readiness levels before a workload is deployed”, the question an acceptance test asks.

Card identity, serial numbers and firmware

nvidia-smi -L lists every GPU in the system with its UUID, and nvidia-smi -q gives the details per card. The serial number in that output “matches the serial number physically printed on each board”, so compare it with the delivery note before the server is racked and the labels are hard to reach. Record the VBIOS version, the Inforom versions and, with -d GSP_FIRMWARE_VERSION, the GSP firmware. Identical cards bought together should report identical versions; raise a card that differs with the supplier before you accept the server.

Install the driver branch the server will run in production before testing; our article on NVIDIA driver branches and CUDA versions covers the choice. Persistence mode set with nvidia-smi -pm “does not persist across reboots”, so start the persistence daemon, nvidia-persistenced, at boot instead.

PCIe link width and speed

A card that trained its link at a lower generation or width still passes a quick check and loses bandwidth until the cause is found. nvidia-smi reports a maximum and a current link generation and width; the maximum is “possible with this GPU and system configuration”, and the current values “may be reduced when the GPU is not in use”. Read them while a test is running, not on an idle server.

NVIDIA lists PCIe Gen5 for the RTX PRO 6000 Server Edition and the H200 NVL, and PCIe Gen4 x16 for the L40S and the L4. Under load, the current values should equal the maximum, and the maximum should match the card unless the server’s slot specification states a lower generation or width. Record “Replays Since Reset” as well and compare it after the burn-in; a rollover after four consecutive replays “results in retraining the link”.

nvidia-smi topo -m shows the connections between all GPUs and NICs and their CPU affinity. Check that each card sits under the processor the configuration planned. On bridged H200 NVL cards, nvidia-smi nvlink -s shows the status of every link and nvidia-smi nvlink -e the error counters. What each slot, riser and cable must be before the card goes in is the subject of our H200 NVL retrofit checklist.

DCGM diagnostics: dcgmi diag levels 1 to 4

dcgmi diag -r takes a run level from 1 to 4, and each level includes the tests of the one below. The software test checks that CUDA applications can run and fails on pages pending retirement and on pending or failed row remaps. The memory test allocates 75 per cent of GPU memory by default and checks written patterns. The PCIe test checks peer-to-peer correctness and allows 80 replays per GPU by default.

LEVELTESTS ADDED4 GPUS8 GPUS
1, shortsoftwareunder 2.5 sunder 2.5 s
2, mediumPCIe and NVLink, GPU memory, memory bandwidthunder 2.5 minunder 10.5 min
3, longdiagnostic, targeted stress, targeted power, nvbandwidth, NCCL testsunder 10 minunder 35 min
4, extra longmemtest, pulse testunder 45 minunder 2.25 h

NVIDIA DCGM diagnostics documentation, version 4.6, read on 10 October 2026; durations as NVIDIA measured them on Hopper GPU systems. NVIDIA’s GPU Debug Guidelines, updated 4 October 2026, say the long test should take approximately 30 minutes.

Add -j for JSON output with a status per test and save it with the server’s records. Exit code 226 means the diagnostic ran and reported an error. NVIDIA describes active health checks as invasive, “requiring exclusive access to the target GPUs”, so run them before any workload is scheduled.

Which levels run depends on the card. DCGM’s feature table uses NVIDIA’s old brand names and gives all levels only to Tesla, the data-centre line, which by our reading covers the H200 NVL, the L40S and the L4; Quadro, Titan and GeForce cards get level 1. The table names no RTX PRO Blackwell card, and the release notes add models by device ID instead. DCGM 4.4.2 added the pulse test, a level 4 test, for “PG153 SKU 210 (devId 2bb5)”, the PCI ID that Canonical’s Ubuntu hardware certification lists for the RTX PRO 6000 Server Edition, and 4.5.3 added the RTX PRO 6000 Max-Q. We found no NVIDIA document that states which run levels pass on the Server Edition. Since 4.6.0, -p "generic_mode=True" lets levels 3 and 4 run on GPUs without SKU-specific calibration, as stability checks without calibrated performance or power thresholds. Run level 3 on the delivered card with the current DCGM, 4.7.0 as of October 2026, and read which tests it reports.

Two level 3 tests also depend on the setup. The NCCL test runs on one node only and needs NCCL and nccl-tests installed, with DCGM_NCCL_TESTS_BIN_PATH pointing to the test binaries before the nvidia-dcgm service starts; without it, DCGM skips the test. The nvbandwidth test runs only on the GPUs DCGM lists for it, among them the L40S, “PG153 SKU 210” and the H200 under device ID 233b, the ID NVIDIA’s driver list gives the H200 NVL. The L4 is not on that list.

Burn-in under sustained power and temperature

For a burn-in, the DCGM test that fits is targeted power. Its goal, in NVIDIA’s words, is “to drive a GPU towards TDP power usage and sustain that throughout the test”. It fails if it cannot reach 75 per cent of the target, and it runs for 120 s by default. NVIDIA’s own example extends it with -p targeted_power.test_duration=600.0, and --iterations repeats a test suite to lengthen the run. Our example, not NVIDIA’s, is a level 3 run followed by several iterations of targeted power at 600 s, so that the chassis fans and the room reach a steady state.

Watch three readings while it runs. The power limit should be the card’s rating unless a lower cap was ordered: up to 600 W on the RTX PRO 6000 Server Edition and the H200 NVL, 350 W on the L40S and 72 W on the L4. A Server Edition card that reports a 450 W limit runs on a cable or in a slot set for 450 W, which some servers do on purpose, so check it against the order, as our guide to servers compatible with the RTX PRO 6000 Server Edition explains. The clocks event reasons in nvidia-smi -q -d PERFORMANCE should show no HW Slowdown or SW Thermal Slowdown; SW Power Cap at full load only means the card is at its power limit. Temperature should stay below the Slowdown Temp the card reports, “the temperature at which a GPU HW will begin optimizing clocks due to thermal conditions”.

Run the burn-in in the rack, at the inlet temperature where the server will work, not on an open bench. Where a DCGM test is skipped or not supported on the card, run the serving engine or training job you will use under a load generator for the same period and watch the same counters.

ECC, row remapping and Xid logs after the run

After the burn-in, nvidia-smi keeps volatile ECC counts “since the last driver load” and aggregate counts that “persist indefinitely”, so compare both with the values recorded before the test. nvidia-smi -q -d ECC,ROW_REMAPPER also shows whether a row remap is pending, which needs a GPU reset to take effect.

Then search the kernel log for Xid messages, for example with journalctl -k | grep -i xid. NVIDIA defines an Xid as “an error report from the NVIDIA driver that is printed to the operating system’s kernel log”. A clean acceptance run should leave none. What each Xid means and which ones need a reset or a drained node is in our article on DCGM metrics and XID errors.

On GPUs we supply, the manufacturer warranty is handled through us and DOA units are replaced. Write to us with the card, its serial number and the diagnostic output if a delivered GPU fails one of these checks.

Multi-GPU bandwidth: nvbandwidth and nccl-tests

nvbandwidth is NVIDIA’s “tool for bandwidth measurements on NVIDIA GPUs”, built from source with cmake . and make. It prints a matrix in GB/s for every GPU pair, with copy engine (CE) and streaming multiprocessor (SM) variants of the tests; -l lists the tests and -t runs one. Every pair of identical cards on identical links should show comparable figures, and a row that falls short points to a link or slot.

nccl-tests check “both the performance and the correctness of NCCL operations”. Build them with make, setting CUDA_HOME and NCCL_HOME if CUDA and NCCL are not in the default paths, then run all_reduce_perf -b 8 -e 128M -f 2 -g 8 from the build directory on an eight-GPU server. The -c option sets how many iterations are checked for correct results, and its default of 1 already checks each run. The busbw column applies the factor 2 × (n − 1) / n to the algorithm bandwidth for all-reduce, so it can be compared “with the hardware peak bandwidth, independently of the number of ranks used”.

Bridged H200 NVL cards exchange data over NVLink at 900 GB/s per GPU in NVIDIA’s figures, against 128 GB/s for PCIe Gen5, both totals for the two directions. A bridge joins at most four cards, so on eight cards the PCIe hops between the NVLink domains limit the all-reduce figure. RTX PRO 6000, L40S and L4 cards have no NVLink and use PCIe.

If PCIe peer-to-peer figures on bare metal come out far below the others, check the PCI bridges. NVIDIA’s NCCL guide says IO virtualisation can redirect peer-to-peer traffic “to the CPU root complex, causing a significant performance reduction or even a hang”, and that ACS might be enabled where the ACSCtl lines of sudo lspci -vvv show SrcValid+. Virtual machines require ACS, so this applies to bare-metal hosts only.

Acceptance procedure, step by step

  1. Before power-on, compare the delivery with the order: chassis model, card count and type, GPU power cables per slot and, if you requested one, the test report.
  2. Install the production driver branch, start nvidia-persistenced at boot, and install DCGM with DCGM_NCCL_TESTS_BIN_PATH set for the nvidia-dcgm service.
  3. Save nvidia-smi -q -x per server; check serials against the labels and firmware versions across identical cards.
  4. Check GPU placement with nvidia-smi topo -m and, on bridged H200 NVL cards, every link with nvidia-smi nvlink -s.
  5. Record ECC counts, row remapper status and PCIe replays as the baseline.
  6. Run dcgmi diag -r 3 -j and keep the JSON; every test should pass.
  7. Run targeted power for an extended period, reading link generation and width, power limit, clocks event reasons and temperatures during the run.
  8. On multi-GPU servers, run nvbandwidth and all_reduce_perf.
  9. Compare ECC counts, remaps and replays with the baseline and search the kernel log for Xid messages. Keep the outputs as the baseline for monitoring later.

Our AI servers are assembled and load-tested before shipment, with a test report on request. Tell us the cards, the number of servers and the checks your acceptance protocol requires in the form below.

What we supply

We build AI servers to order around the workload, assembled and burn-in tested. Operating system, drivers, CUDA and a container runtime are installed on request. We also supply the professional NVIDIA GPUs these checks cover, the RTX PRO 6000 Server Edition, the H200 NVL with its NVLink bridges, the L40S and the L4, for servers you already run, with manufacturer warranty on one EU contract and invoice. We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day.

FAQ

How do you burn in or stress test a GPU server?
Run a sustained load that keeps every GPU near its power limit while you watch power, clocks event reasons and temperatures, then check ECC counters, row remaps and the kernel log for Xid messages. DCGM’s targeted power test drives each GPU towards its TDP and can be extended with a test_duration parameter and the --iterations option. NVIDIA publishes no burn-in duration, so the length of the run is your own acceptance criterion.
What does dcgmi diag level 3 test?
Level 3 runs the software, PCIe and NVLink, GPU memory and memory bandwidth tests of level 2 and adds the diagnostic, targeted stress, targeted power, nvbandwidth and NCCL tests. The DCGM documentation gives its duration as under 10 minutes on 4-GPU and under 35 minutes on 8-GPU Hopper systems, while NVIDIA’s GPU triage guide says the long test should take approximately 30 minutes. The NCCL test runs only if DCGM_NCCL_TESTS_BIN_PATH is set for the DCGM service, and the nvbandwidth test only on the GPUs DCGM lists for it.
Do DCGM diagnostics work on RTX PRO cards?
DCGM’s feature table gives all diagnostic levels to Tesla, NVIDIA’s data-centre line, and only level 1 to Quadro, Titan and GeForce cards, and it names no RTX PRO Blackwell card. The release notes add RTX PRO models by device ID, and since DCGM 4.6.0 levels 3 and 4 can run on GPUs without SKU-specific calibration with the generic_mode=True parameter, as stability checks. Run dcgmi diag -r 3 with the current DCGM on the delivered card and read which tests it reports.
What should a GPU server acceptance test include?
It should confirm that every ordered card is present with the expected serial numbers, firmware and PCIe link, that DCGM diagnostics pass, that the GPUs hold their power limit without thermal slowdown under sustained load, and that no ECC errors, pending row remaps or Xid messages appear afterwards. On multi-GPU servers, add nvbandwidth and nccl-tests. Save all outputs as the baseline for later monitoring.
How do I run nccl-tests on a GPU server?
Build the tests with make, then run ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8 on an eight-GPU server; the results are checked for correctness by default. Compare the busbw column with the peak bandwidth of the link, NVLink on bridged H200 NVL cards and PCIe on the other cards, keeping in mind that NVIDIA’s 900 and 128 GB/s are totals for both directions. Low peer-to-peer figures on bare metal can point to PCI ACS redirecting traffic through the CPU root complex.
Why does nvidia-smi show a lower PCIe generation than expected?
nvidia-smi notes that the current link generation and width may be reduced when the GPU is not in use, so an idle card can show a lower value. Read the link while a test is running and compare it with the maximum nvidia-smi reports. NVIDIA lists Gen5 for the RTX PRO 6000 Server Edition and H200 NVL and Gen4 x16 for the L40S and L4.

Send us the cards, the number of servers, the server models and the checks your acceptance protocol requires. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna