XID 79 “GPU has fallen off the bus”: causes, checks in order and when to open a warranty case
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- XID 79 means the NVIDIA driver tried to reach a GPU over PCI Express and found it not accessible; NVIDIA’s Xid catalogue gives a system restart (RESTART_BM) as the immediate action and contacting support as the investigatory action
- NVIDIA names the cause classes as hardware failures on the PCIe link, failing GPU hardware and driver issues, and points to the system event log and the kernel’s PCI event logs for the source of the link failure
- Collect nvidia-bug-report.sh output, the kernel log and the BMC event log before the restart; NVIDIA’s debug guidelines say to drain the node and report the issue to the system vendor
- Check in order: link and replays after the restart, power cable and supplies, thermals, seating and risers, firmware, driver; then swap slots to tell a faulty card from a faulty slot or platform
- NVIDIA calls its field diagnostic the authoritative tool for GPU health and says it is usually required before an RMA can start; the system vendor says when and how to run it
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What XID 79 “GPU has fallen off the bus” means
XID 79 is the NVIDIA driver’s report that it tried to reach a GPU over its PCI Express connection and found the GPU not accessible. NVIDIA’s Xid catalogue (release 615, 9 September 2026) gives RESTART_BM as the immediate action, defined as “Restart bare metal, system should be restarted”, and CONTACT_SUPPORT as the investigatory action. In the kernel log the line starts with NVRM: Xid, carries the PCI address of the affected card, the code 79 and the text “GPU has fallen off the bus”.
The catalogue’s trigger conditions name the cause classes. NVIDIA writes that the event “is often caused by hardware failures on the PCI Express link”, that failing GPU hardware or driver issues may also cause it, and that system event logs and kernel PCI event logs may show where the link failure came from. NVIDIA’s GPU debug guidelines (updated 4 October 2026) add the operational step for 79, which is to drain the node and report the issue. The catalogue also lists 79 among the codes that trigger Xid 154 from driver branch R565 on, a second message that states the recovery action, for example “Node Reboot Required”.
The other codes are in our NVIDIA XID error codes reference; this article covers 79 alone.
First response: drain the node, collect logs, then restart
Take the node out of the load balancer or the scheduler first, because the catalogue’s action for 79 is a restart of the whole system. NVIDIA’s management library returns NVML_ERROR_GPU_IS_LOST for such a card, described as “The GPU has fallen off the bus or has otherwise become inaccessible”. nvidia-smi draws much of its function from that library, so the card fails or is missing in its output until the restart.
Collect the evidence before the restart, because a reboot clears the kernel’s message buffer and, on hosts that keep the kernel log only in a volatile journal, the log with it. NVIDIA says nvidia-bug-report.sh, run as root, installs with the driver, writes one file, nvidia-bug-report.log.gz, may take up to an hour, and can be run with --safe-mode --extra-system-data if it hangs. Export the BMC event log as well, note the time of the XID line and check the BMC clock against the host’s, so the two logs can be lined up.
A GPU reset with nvidia-smi -r can clear GPU hardware and software state “in situations that would otherwise require a machine reboot”. It requires root and no applications on the device, and NVIDIA states that a reset does not succeed in every case and that “It is not recommended for production environments at this time.” For 79 the catalogue asks for a system restart rather than a GPU reset. If the card is still missing after a warm reboot, power the server off fully; for a GPU that stays unhealthy, the nvidia-smi documentation says “a complete reset should be instigated by power cycling the node”. For alerts on 79, see our article on DCGM metrics and XID errors.
Where the evidence is: kernel log, PCIe AER and the BMC
NVIDIA’s Xid documentation says to grep for “NVRM: Xid” in /var/log/messages or /var/log/syslog, and journalctl -k shows the kernel messages of the current boot, including the lines before the XID. Where the Linux PCIe AER driver handles errors, it logs them as “PCIe Bus Error” lines with a severity. The kernel documentation separates correctable errors, which “pose no impacts on the functionality of the interface”, from non-fatal errors, which leave the link working, and fatal errors, which “cause the link to be unreliable”.
The per-device counters are in sysfs, in the files aer_dev_correctable, aer_dev_nonfatal and aer_dev_fatal under /sys/bus/pci/devices/ and the card’s address. Both uncorrectable files list error types such as Surprise Down Error and Completion Timeout, each with its own count. Read them on the GPU and on the port above it. Linux handles AER only when the firmware grants it control through ACPI _OSC, and on platforms where it does not, look for the errors in the BMC event log.
| SYMPTOM | CHECK | SOURCE |
|---|---|---|
| Card back after restart | link generation, width and “Replays Since Reset” under load in nvidia-smi -q | nvidia-smi documentation |
| AER fatal error before XID | seating, riser and riser cable; aer_ | Linux kernel AER documentation |
| PCIe error in BMC log | server maker’s KB and firmware for BIOS, BMC and GPU | NVIDIA Xid catalogue |
| HW Power Brake | number and rating of power supplies, GPU cable for that slot | nvidia-smi; maker’s GPU rules |
| HW Thermal Slowdown | fan modules, air duct, inlet temperature | nvidia-smi; maker’s thermal rules |
| Card absent from first boot | power class the 16-pin cable declares, cable part number per slot | NVIDIA H200 NVL product brief |
NVIDIA Xid catalogue and nvidia-smi documentation, Linux kernel PCIe AER documentation and NVIDIA H200 NVL product brief PB-12128-001_v01, read on 10 October 2026.
Troubleshooting order for XID 79
Write down the result of each step for the case.
- Record the XID line, the PCI address, the time and the workload, and collect nvidia-bug-report.sh output and the BMC event log before any restart.
- Restart the host, with a full power-off if the card stays missing, then read the link generation, width and replays under load with nvidia-smi -q.
- Check power, meaning the GPU cable’s part number for that slot, the power supplies the server maker requires for your GPU count and the HW Power Brake clock event reason under load.
- Check thermals against the server maker’s rules for fan modules, air duct and inlet temperature, and watch the HW Thermal Slowdown reason under load.
- Reseat the card and check the riser and its cables; the catalogue’s own wording for mechanical checks is to ensure “that device seating and all applicable connections to it are secure”.
- Compare BIOS, BMC, GPU and, where the server lists it, PCIe retimer firmware with the server maker’s current releases and search its knowledge base for your model.
- Confirm that the driver branch supports the card and read the known issues in that branch’s release notes.
- If 79 returns, swap slots and run
dcgmi diag -r 3on the affected GPU before you open the case.
nvidia-smi notes that the current link values “may be reduced when the GPU is not in use”, so read them under load, and that a replay number rollover after 4 consecutive replays “results in retraining the link”. The release notes of driver 580.178.04 (October 2026) list no XID 79 issue among the fixed or known issues.
Power and thermals in 4- to 8-card servers
NVIDIA’s catalogue does not name power or temperature for 79. The order above checks both early, because server makers support a GPU configuration only under stated power and cooling conditions.
With 600 W cards, check the power cable first. The H200 NVL product brief (PB-12128-001_v01) marks every power level of the 16-pin cable of 450 W or less “Not supported. Insufficient power”, and an H200 NVL whose cable declares less than its default power cap “will not boot”. That failure shows as a card missing from the first power-on, not as an XID 79 on a card that has been running. NVIDIA’s RTX PRO 6000 Blackwell Server Edition brief (SP-12355-001_v02) lists a 600 W and a 450 W mode. Lenovo Press LP2128 lists the SR650a V4 with four RTX PRO 6000 Server Edition only with the slot capped at 450 W, and with two of them or two H200 NVL at 600 W. In a 4-card server, check that the cable and power mode match what the maker specifies for the slot.
Eight 600 W cards draw 4.8 kW on their own, and server makers set the number and rating of power supplies per GPU count, as our comparison of how many GPUs fit in one server shows. In nvidia-smi, the clock event reason HW Power Brake means “External Power Brake Assertion is triggered (e.g. by the system power supply)”. Read it with nvidia-smi -q -d PERFORMANCE while every card runs at full load after the restart.
Passive cards take their cooling from the chassis. NVIDIA writes that the H200 NVL’s heat sink “requires system airflow to operate the card properly”, and the brief allows 10 to 45 °C ambient for the card. Server makers set lower inlet limits per model, slot and fan module, as our H200 NVL retrofit checklist shows. HW Thermal Slowdown reports clocks reduced “by a factor of 2 or more due to temperature being too high”. If the BMC or your monitoring records GPU temperatures out of band, read the minutes before the XID.
When we build on a chassis you already own, we check its platform, power and cooling first and say in advance if something will not work. Tell us the server model, its slots and the cards you plan in it.
One faulty card or a slot problem: the swap test
NVIDIA describes the investigatory action as guidance for an issue that recurs or is unexpected, so a second XID 79 starts the isolation. Note each card’s serial number and bus address from nvidia-smi -q first, because the address belongs to the slot and the serial to the card.
Move the affected card to a slot whose GPU has run without errors, and that GPU into the suspect slot, keeping the driver and the workload unchanged. Use only slots the server maker permits for that card, with the power cable that slot requires. If the next XID follows the serial number, the card is the suspect. If it stays with the slot, look at the riser, the slot’s cable and the PCIe switch or retimer in that path. If XIDs appear on different GPUs in turn, look at the platform, meaning firmware, power supplies and switch boards.
Run DCGM diagnostics on the moved card with dcgmi diag -r 3 -i 2, where 2 is its GPU ID. The DCGM documentation gives level 3 under 10 minutes on 4-GPU and under 35 minutes on 8-GPU Hopper systems; NVIDIA’s triage guide says the long test should take approximately 30 minutes. The PCIe test allows 80 replays per GPU by default. Which levels run on RTX PRO cards is covered in our acceptance testing guide.
Take, as an example, a server with eight RTX PRO 6000 Server Edition cards, four behind each PCIe switch, that logs 79 twice for one bus address. A third XID 79 at the new address under the same serial points at the card. XIDs on two cards behind one switch, with a bus fatal error in the BMC log, point at the switch path, and the server maker’s support comes first.
What a warranty case for XID 79 needs
Open the case once 79 returns after the checks above, when the catalogue’s investigatory action, contacting support, applies. Who handles it, the card’s supplier or the server maker, and what evidence it accepts follow that party’s warranty terms. NVIDIA’s debug guidelines list what a GPU issue report should contain and end with “Submit a ticket to your system vendor”. They also describe NVIDIA’s field diagnostic as “the authoritative and comprehensive NVIDIA tool for determining the health of GPUs”, which “is usually required before an RMA can be started”. NVIDIA leaves when and how to run it to the system vendor.
| EVIDENCE | HOW TO COLLECT | ASKED FOR BY |
|---|---|---|
| OS and driver details | nvidia-bug-report. | NVIDIA debug guidelines |
| Key log lines, full log | kernel log with XID and AER lines | NVIDIA debug guidelines |
| Debug steps taken | the troubleshooting order above, with results | NVIDIA debug guidelines |
| DCGM diagnostics log | dcgmi diag -r 3 with -j for JSON | NVIDIA debug guidelines |
| BMC event log | export from the BMC, time of the XID noted | our recommendation |
| Serial number and slot | nvidia-smi -q and the label on the card | our recommendation |
| Field diagnostic | run only on the vendor’s instruction | NVIDIA debug guidelines |
NVIDIA GPU debug guidelines, GPU node triage, “Reporting a GPU Issue” and “Running Field Diagnostics”, updated 4 October 2026; rows marked «our recommendation» are ours.
For GPUs we supply, the manufacturer warranty is handled through us and support by our engineering partner’s service team. Send us the card, its serial number and the XID lines through the form below.
What we supply
We supply the cards this article discusses, the RTX PRO 6000 Server Edition and the H200 NVL with its NVLink bridges, and build AI servers to order around them, assembled and burn-in tested, with manufacturer warranty on every component. For the professional NVIDIA GPUs we supply, the manufacturer warranty is handled through us, DOA units are replaced, and support is handled by our engineering partner’s service team. When we build on your own chassis or with parts you already own, we check the platform, power and cooling first and say in advance if something will not work. Configuration and quote follow within one business day, on one EU contract and invoice.
FAQ
What does XID 79 mean?
How do I fix XID 79 GPU has fallen off the bus?
Why does nvidia-smi show no devices after XID 79?
Can I reset the GPU instead of rebooting after XID 79?
Is XID 79 a hardware failure?
What do I need for an RMA after XID 79?
Send us the server model, the GPU cards and the slots they sit in, the XID lines with their PCI addresses and the power supplies fitted. We reply within one business day; for cards bought from us, the manufacturer warranty is handled through us.
Talk to an expertWe reply within one business day