GPU server acceptance testing: burn-in, DCGM diagnostics and what to check on delivery
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Accept a GPU server in four steps: identity (every card, its serial, VBIOS and PCIe link), health with DCGM diagnostics, power and temperature under sustained load, and GPU-to-GPU bandwidth where the cards exchange data
- DCGM diagnostics have four run levels; level 3 adds the diagnostic, targeted stress, targeted power, nvbandwidth and NCCL tests to level 2 and takes under 10 minutes on a 4-GPU Hopper system by the DCGM documentation, while NVIDIA’s triage guide says the long test should take approximately 30 minutes
- DCGM’s feature table gives every diagnostic level only to Tesla, NVIDIA’s old name for its data-centre line, and level 1 to Quadro, Titan and GeForce; since DCGM 4.6.0, levels 3 and 4 can run as stability checks on GPUs without SKU-specific calibration
- Read the PCIe link while the GPU is busy, because nvidia-smi says the current generation and width may be reduced when the GPU is not in use; NVIDIA lists Gen5 for the RTX PRO 6000 Server Edition and H200 NVL and Gen4 x16 for the L40S and L4
- Targeted power drives each GPU towards its TDP and fails below 75 per cent of the target by default; as our acceptance criterion, uncorrectable ECC errors, pending row remaps and Xid messages after the burn-in should all be zero
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What to check when a GPU server is delivered
A GPU server acceptance test, with a burn-in, checks four things before the server goes into production. First confirm its identity, meaning that every ordered card is present with the expected serial number, VBIOS and PCIe link. Then check its health with NVIDIA’s DCGM diagnostics, its power and temperature under sustained load and, on servers where the cards exchange data, the bandwidth between GPUs. The table lists each check, its tool and the pass criterion.
| CHECK | TOOL | PASS CRITERION | BASIS |
|---|---|---|---|
| Cards and serials | nvidia-smi -L, nvidia-smi -q | every ordered card listed; serials match the board labels | NVIDIA |
| VBIOS and firmware | nvidia-smi -q | identical versions on identical cards | our criterion |
| PCIe link | nvidia-smi -q under load | current generation and width equal the maximum | NVIDIA, link drops when idle |
| Health | dcgmi diag -r 3 | every test passes | NVIDIA DCGM |
| Sustained power | dcgmi diag -r targeted_ | at least 75 per cent of target power | DCGM default |
| Clocks and temperature | nvidia-smi -q -d PERFORMANCE | no thermal slowdown during the run | our criterion |
| Memory errors | nvidia-smi -q -d ECC,ROW_ | no uncorrectable errors, no pending remaps | DCGM software test |
| Driver errors | kernel log | no Xid messages during the run | our criterion |
| GPU-to-GPU bandwidth | nvbandwidth, nccl-tests | comparable figures on every GPU pair | our criterion |
nvidia-smi documentation and DCGM 4.6 diagnostics documentation, read on 10 October 2026; rows marked «our criterion» are our examples, not NVIDIA pass thresholds.
We found no NVIDIA document that sets a burn-in duration or calls DCGM an acceptance test. DCGM’s documentation lists among its goals a tool “to assess cluster readiness levels before a workload is deployed”, the question an acceptance test asks.
Card identity, serial numbers and firmware
nvidia-smi -L lists every GPU in the system with its UUID, and nvidia-smi -q gives the details per card. The serial number in that output “matches the serial number physically printed on each board”, so compare it with the delivery note before the server is racked and the labels are hard to reach. Record the VBIOS version, the Inforom versions and, with -d GSP_FIRMWARE_VERSION, the GSP firmware. Identical cards bought together should report identical versions; raise a card that differs with the supplier before you accept the server.
Install the driver branch the server will run in production before testing; our article on NVIDIA driver branches and CUDA versions covers the choice. Persistence mode set with nvidia-smi -pm “does not persist across reboots”, so start the persistence daemon, nvidia-persistenced, at boot instead.
PCIe link width and speed
A card that trained its link at a lower generation or width still passes a quick check and loses bandwidth until the cause is found. nvidia-smi reports a maximum and a current link generation and width; the maximum is “possible with this GPU and system configuration”, and the current values “may be reduced when the GPU is not in use”. Read them while a test is running, not on an idle server.
NVIDIA lists PCIe Gen5 for the RTX PRO 6000 Server Edition and the H200 NVL, and PCIe Gen4 x16 for the L40S and the L4. Under load, the current values should equal the maximum, and the maximum should match the card unless the server’s slot specification states a lower generation or width. Record “Replays Since Reset” as well and compare it after the burn-in; a rollover after four consecutive replays “results in retraining the link”.
nvidia-smi topo -m shows the connections between all GPUs and NICs and their CPU affinity. Check that each card sits under the processor the configuration planned. On bridged H200 NVL cards, nvidia-smi nvlink -s shows the status of every link and nvidia-smi nvlink -e the error counters. What each slot, riser and cable must be before the card goes in is the subject of our H200 NVL retrofit checklist.
DCGM diagnostics: dcgmi diag levels 1 to 4
dcgmi diag -r takes a run level from 1 to 4, and each level includes the tests of the one below. The software test checks that CUDA applications can run and fails on pages pending retirement and on pending or failed row remaps. The memory test allocates 75 per cent of GPU memory by default and checks written patterns. The PCIe test checks peer-to-peer correctness and allows 80 replays per GPU by default.
| LEVEL | TESTS ADDED | 4 GPUS | 8 GPUS |
|---|---|---|---|
| 1, short | software | under 2.5 s | under 2.5 s |
| 2, medium | PCIe and NVLink, GPU memory, memory bandwidth | under 2.5 min | under 10.5 min |
| 3, long | diagnostic, targeted stress, targeted power, nvbandwidth, NCCL tests | under 10 min | under 35 min |
| 4, extra long | memtest, pulse test | under 45 min | under 2.25 h |
NVIDIA DCGM diagnostics documentation, version 4.6, read on 10 October 2026; durations as NVIDIA measured them on Hopper GPU systems. NVIDIA’s GPU Debug Guidelines, updated 4 October 2026, say the long test should take approximately 30 minutes.
Add -j for JSON output with a status per test and save it with the server’s records. Exit code 226 means the diagnostic ran and reported an error. NVIDIA describes active health checks as invasive, “requiring exclusive access to the target GPUs”, so run them before any workload is scheduled.
Which levels run depends on the card. DCGM’s feature table uses NVIDIA’s old brand names and gives all levels only to Tesla, the data-centre line, which by our reading covers the H200 NVL, the L40S and the L4; Quadro, Titan and GeForce cards get level 1. The table names no RTX PRO Blackwell card, and the release notes add models by device ID instead. DCGM 4.4.2 added the pulse test, a level 4 test, for “PG153 SKU 210 (devId 2bb5)”, the PCI ID that Canonical’s Ubuntu hardware certification lists for the RTX PRO 6000 Server Edition, and 4.5.3 added the RTX PRO 6000 Max-Q. We found no NVIDIA document that states which run levels pass on the Server Edition. Since 4.6.0, -p "generic_mode=True" lets levels 3 and 4 run on GPUs without SKU-specific calibration, as stability checks without calibrated performance or power thresholds. Run level 3 on the delivered card with the current DCGM, 4.7.0 as of October 2026, and read which tests it reports.
Two level 3 tests also depend on the setup. The NCCL test runs on one node only and needs NCCL and nccl-tests installed, with DCGM_NCCL_TESTS_BIN_PATH pointing to the test binaries before the nvidia-dcgm service starts; without it, DCGM skips the test. The nvbandwidth test runs only on the GPUs DCGM lists for it, among them the L40S, “PG153 SKU 210” and the H200 under device ID 233b, the ID NVIDIA’s driver list gives the H200 NVL. The L4 is not on that list.
Burn-in under sustained power and temperature
For a burn-in, the DCGM test that fits is targeted power. Its goal, in NVIDIA’s words, is “to drive a GPU towards TDP power usage and sustain that throughout the test”. It fails if it cannot reach 75 per cent of the target, and it runs for 120 s by default. NVIDIA’s own example extends it with -p targeted_, and --iterations repeats a test suite to lengthen the run. Our example, not NVIDIA’s, is a level 3 run followed by several iterations of targeted power at 600 s, so that the chassis fans and the room reach a steady state.
Watch three readings while it runs. The power limit should be the card’s rating unless a lower cap was ordered: up to 600 W on the RTX PRO 6000 Server Edition and the H200 NVL, 350 W on the L40S and 72 W on the L4. A Server Edition card that reports a 450 W limit runs on a cable or in a slot set for 450 W, which some servers do on purpose, so check it against the order, as our guide to servers compatible with the RTX PRO 6000 Server Edition explains. The clocks event reasons in nvidia-smi -q -d PERFORMANCE should show no HW Slowdown or SW Thermal Slowdown; SW Power Cap at full load only means the card is at its power limit. Temperature should stay below the Slowdown Temp the card reports, “the temperature at which a GPU HW will begin optimizing clocks due to thermal conditions”.
Run the burn-in in the rack, at the inlet temperature where the server will work, not on an open bench. Where a DCGM test is skipped or not supported on the card, run the serving engine or training job you will use under a load generator for the same period and watch the same counters.
ECC, row remapping and Xid logs after the run
After the burn-in, nvidia-smi keeps volatile ECC counts “since the last driver load” and aggregate counts that “persist indefinitely”, so compare both with the values recorded before the test. nvidia-smi -q -d ECC,ROW_ also shows whether a row remap is pending, which needs a GPU reset to take effect.
Then search the kernel log for Xid messages, for example with journalctl -k | grep -i xid. NVIDIA defines an Xid as “an error report from the NVIDIA driver that is printed to the operating system’s kernel log”. A clean acceptance run should leave none. What each Xid means and which ones need a reset or a drained node is in our article on DCGM metrics and XID errors.
On GPUs we supply, the manufacturer warranty is handled through us and DOA units are replaced. Write to us with the card, its serial number and the diagnostic output if a delivered GPU fails one of these checks.
Multi-GPU bandwidth: nvbandwidth and nccl-tests
nvbandwidth is NVIDIA’s “tool for bandwidth measurements on NVIDIA GPUs”, built from source with cmake . and make. It prints a matrix in GB/s for every GPU pair, with copy engine (CE) and streaming multiprocessor (SM) variants of the tests; -l lists the tests and -t runs one. Every pair of identical cards on identical links should show comparable figures, and a row that falls short points to a link or slot.
nccl-tests check “both the performance and the correctness of NCCL operations”. Build them with make, setting CUDA_HOME and NCCL_HOME if CUDA and NCCL are not in the default paths, then run all_ from the build directory on an eight-GPU server. The -c option sets how many iterations are checked for correct results, and its default of 1 already checks each run. The busbw column applies the factor 2 × (n − 1) / n to the algorithm bandwidth for all-reduce, so it can be compared “with the hardware peak bandwidth, independently of the number of ranks used”.
Bridged H200 NVL cards exchange data over NVLink at 900 GB/s per GPU in NVIDIA’s figures, against 128 GB/s for PCIe Gen5, both totals for the two directions. A bridge joins at most four cards, so on eight cards the PCIe hops between the NVLink domains limit the all-reduce figure. RTX PRO 6000, L40S and L4 cards have no NVLink and use PCIe.
If PCIe peer-to-peer figures on bare metal come out far below the others, check the PCI bridges. NVIDIA’s NCCL guide says IO virtualisation can redirect peer-to-peer traffic “to the CPU root complex, causing a significant performance reduction or even a hang”, and that ACS might be enabled where the ACSCtl lines of sudo lspci -vvv show SrcValid+. Virtual machines require ACS, so this applies to bare-metal hosts only.
Acceptance procedure, step by step
- Before power-on, compare the delivery with the order: chassis model, card count and type, GPU power cables per slot and, if you requested one, the test report.
- Install the production driver branch, start nvidia-persistenced at boot, and install DCGM with
DCGM_NCCL_TESTS_BIN_PATHset for the nvidia-dcgm service. - Save
nvidia-smi -q -xper server; check serials against the labels and firmware versions across identical cards. - Check GPU placement with
nvidia-smi topo -mand, on bridged H200 NVL cards, every link withnvidia-smi nvlink -s. - Record ECC counts, row remapper status and PCIe replays as the baseline.
- Run
dcgmi diag -r 3 -jand keep the JSON; every test should pass. - Run targeted power for an extended period, reading link generation and width, power limit, clocks event reasons and temperatures during the run.
- On multi-GPU servers, run nvbandwidth and
all_reduce_perf. - Compare ECC counts, remaps and replays with the baseline and search the kernel log for Xid messages. Keep the outputs as the baseline for monitoring later.
Our AI servers are assembled and load-tested before shipment, with a test report on request. Tell us the cards, the number of servers and the checks your acceptance protocol requires in the form below.
What we supply
We build AI servers to order around the workload, assembled and burn-in tested. Operating system, drivers, CUDA and a container runtime are installed on request. We also supply the professional NVIDIA GPUs these checks cover, the RTX PRO 6000 Server Edition, the H200 NVL with its NVLink bridges, the L40S and the L4, for servers you already run, with manufacturer warranty on one EU contract and invoice. We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day.
FAQ
How do you burn in or stress test a GPU server?
What does dcgmi diag level 3 test?
Do DCGM diagnostics work on RTX PRO cards?
What should a GPU server acceptance test include?
How do I run nccl-tests on a GPU server?
Why does nvidia-smi show a lower PCIe generation than expected?
Send us the cards, the number of servers, the server models and the checks your acceptance protocol requires. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day