BLOG · GUIDE ·

XID 79 “GPU has fallen off the bus”: causes, checks in order and when to open a warranty case

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • XID 79 means the NVIDIA driver tried to reach a GPU over PCI Express and found it not accessible; NVIDIA’s Xid catalogue gives a system restart (RESTART_BM) as the immediate action and contacting support as the investigatory action
  • NVIDIA names the cause classes as hardware failures on the PCIe link, failing GPU hardware and driver issues, and points to the system event log and the kernel’s PCI event logs for the source of the link failure
  • Collect nvidia-bug-report.sh output, the kernel log and the BMC event log before the restart; NVIDIA’s debug guidelines say to drain the node and report the issue to the system vendor
  • Check in order: link and replays after the restart, power cable and supplies, thermals, seating and risers, firmware, driver; then swap slots to tell a faulty card from a faulty slot or platform
  • NVIDIA calls its field diagnostic the authoritative tool for GPU health and says it is usually required before an RMA can start; the system vendor says when and how to run it

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What XID 79 “GPU has fallen off the bus” means

XID 79 is the NVIDIA driver’s report that it tried to reach a GPU over its PCI Express connection and found the GPU not accessible. NVIDIA’s Xid catalogue (release 615, 9 September 2026) gives RESTART_BM as the immediate action, defined as “Restart bare metal, system should be restarted”, and CONTACT_SUPPORT as the investigatory action. In the kernel log the line starts with NVRM: Xid, carries the PCI address of the affected card, the code 79 and the text “GPU has fallen off the bus”.

The catalogue’s trigger conditions name the cause classes. NVIDIA writes that the event “is often caused by hardware failures on the PCI Express link”, that failing GPU hardware or driver issues may also cause it, and that system event logs and kernel PCI event logs may show where the link failure came from. NVIDIA’s GPU debug guidelines (updated 4 October 2026) add the operational step for 79, which is to drain the node and report the issue. The catalogue also lists 79 among the codes that trigger Xid 154 from driver branch R565 on, a second message that states the recovery action, for example “Node Reboot Required”.

The other codes are in our NVIDIA XID error codes reference; this article covers 79 alone.

First response: drain the node, collect logs, then restart

Take the node out of the load balancer or the scheduler first, because the catalogue’s action for 79 is a restart of the whole system. NVIDIA’s management library returns NVML_ERROR_GPU_IS_LOST for such a card, described as “The GPU has fallen off the bus or has otherwise become inaccessible”. nvidia-smi draws much of its function from that library, so the card fails or is missing in its output until the restart.

Collect the evidence before the restart, because a reboot clears the kernel’s message buffer and, on hosts that keep the kernel log only in a volatile journal, the log with it. NVIDIA says nvidia-bug-report.sh, run as root, installs with the driver, writes one file, nvidia-bug-report.log.gz, may take up to an hour, and can be run with --safe-mode --extra-system-data if it hangs. Export the BMC event log as well, note the time of the XID line and check the BMC clock against the host’s, so the two logs can be lined up.

A GPU reset with nvidia-smi -r can clear GPU hardware and software state “in situations that would otherwise require a machine reboot”. It requires root and no applications on the device, and NVIDIA states that a reset does not succeed in every case and that “It is not recommended for production environments at this time.” For 79 the catalogue asks for a system restart rather than a GPU reset. If the card is still missing after a warm reboot, power the server off fully; for a GPU that stays unhealthy, the nvidia-smi documentation says “a complete reset should be instigated by power cycling the node”. For alerts on 79, see our article on DCGM metrics and XID errors.

Where the evidence is: kernel log, PCIe AER and the BMC

NVIDIA’s Xid documentation says to grep for “NVRM: Xid” in /var/log/messages or /var/log/syslog, and journalctl -k shows the kernel messages of the current boot, including the lines before the XID. Where the Linux PCIe AER driver handles errors, it logs them as “PCIe Bus Error” lines with a severity. The kernel documentation separates correctable errors, which “pose no impacts on the functionality of the interface”, from non-fatal errors, which leave the link working, and fatal errors, which “cause the link to be unreliable”.

The per-device counters are in sysfs, in the files aer_dev_correctable, aer_dev_nonfatal and aer_dev_fatal under /sys/bus/pci/devices/ and the card’s address. Both uncorrectable files list error types such as Surprise Down Error and Completion Timeout, each with its own count. Read them on the GPU and on the port above it. Linux handles AER only when the firmware grants it control through ACPI _OSC, and on platforms where it does not, look for the errors in the BMC event log.

SYMPTOMCHECKSOURCE
Card back after restartlink generation, width and “Replays Since Reset” under load in nvidia-smi -qnvidia-smi documentation
AER fatal error before XIDseating, riser and riser cable; aer_dev_fatal on GPU and upstream portLinux kernel AER documentation
PCIe error in BMC logserver maker’s KB and firmware for BIOS, BMC and GPUNVIDIA Xid catalogue
HW Power Brakenumber and rating of power supplies, GPU cable for that slotnvidia-smi; maker’s GPU rules
HW Thermal Slowdownfan modules, air duct, inlet temperaturenvidia-smi; maker’s thermal rules
Card absent from first bootpower class the 16-pin cable declares, cable part number per slotNVIDIA H200 NVL product brief

NVIDIA Xid catalogue and nvidia-smi documentation, Linux kernel PCIe AER documentation and NVIDIA H200 NVL product brief PB-12128-001_v01, read on 10 October 2026.

Troubleshooting order for XID 79

Write down the result of each step for the case.

  1. Record the XID line, the PCI address, the time and the workload, and collect nvidia-bug-report.sh output and the BMC event log before any restart.
  2. Restart the host, with a full power-off if the card stays missing, then read the link generation, width and replays under load with nvidia-smi -q.
  3. Check power, meaning the GPU cable’s part number for that slot, the power supplies the server maker requires for your GPU count and the HW Power Brake clock event reason under load.
  4. Check thermals against the server maker’s rules for fan modules, air duct and inlet temperature, and watch the HW Thermal Slowdown reason under load.
  5. Reseat the card and check the riser and its cables; the catalogue’s own wording for mechanical checks is to ensure “that device seating and all applicable connections to it are secure”.
  6. Compare BIOS, BMC, GPU and, where the server lists it, PCIe retimer firmware with the server maker’s current releases and search its knowledge base for your model.
  7. Confirm that the driver branch supports the card and read the known issues in that branch’s release notes.
  8. If 79 returns, swap slots and run dcgmi diag -r 3 on the affected GPU before you open the case.

nvidia-smi notes that the current link values “may be reduced when the GPU is not in use”, so read them under load, and that a replay number rollover after 4 consecutive replays “results in retraining the link”. The release notes of driver 580.178.04 (October 2026) list no XID 79 issue among the fixed or known issues.

Power and thermals in 4- to 8-card servers

NVIDIA’s catalogue does not name power or temperature for 79. The order above checks both early, because server makers support a GPU configuration only under stated power and cooling conditions.

With 600 W cards, check the power cable first. The H200 NVL product brief (PB-12128-001_v01) marks every power level of the 16-pin cable of 450 W or less “Not supported. Insufficient power”, and an H200 NVL whose cable declares less than its default power cap “will not boot”. That failure shows as a card missing from the first power-on, not as an XID 79 on a card that has been running. NVIDIA’s RTX PRO 6000 Blackwell Server Edition brief (SP-12355-001_v02) lists a 600 W and a 450 W mode. Lenovo Press LP2128 lists the SR650a V4 with four RTX PRO 6000 Server Edition only with the slot capped at 450 W, and with two of them or two H200 NVL at 600 W. In a 4-card server, check that the cable and power mode match what the maker specifies for the slot.

Eight 600 W cards draw 4.8 kW on their own, and server makers set the number and rating of power supplies per GPU count, as our comparison of how many GPUs fit in one server shows. In nvidia-smi, the clock event reason HW Power Brake means “External Power Brake Assertion is triggered (e.g. by the system power supply)”. Read it with nvidia-smi -q -d PERFORMANCE while every card runs at full load after the restart.

Passive cards take their cooling from the chassis. NVIDIA writes that the H200 NVL’s heat sink “requires system airflow to operate the card properly”, and the brief allows 10 to 45 °C ambient for the card. Server makers set lower inlet limits per model, slot and fan module, as our H200 NVL retrofit checklist shows. HW Thermal Slowdown reports clocks reduced “by a factor of 2 or more due to temperature being too high”. If the BMC or your monitoring records GPU temperatures out of band, read the minutes before the XID.

When we build on a chassis you already own, we check its platform, power and cooling first and say in advance if something will not work. Tell us the server model, its slots and the cards you plan in it.

One faulty card or a slot problem: the swap test

NVIDIA describes the investigatory action as guidance for an issue that recurs or is unexpected, so a second XID 79 starts the isolation. Note each card’s serial number and bus address from nvidia-smi -q first, because the address belongs to the slot and the serial to the card.

Move the affected card to a slot whose GPU has run without errors, and that GPU into the suspect slot, keeping the driver and the workload unchanged. Use only slots the server maker permits for that card, with the power cable that slot requires. If the next XID follows the serial number, the card is the suspect. If it stays with the slot, look at the riser, the slot’s cable and the PCIe switch or retimer in that path. If XIDs appear on different GPUs in turn, look at the platform, meaning firmware, power supplies and switch boards.

Run DCGM diagnostics on the moved card with dcgmi diag -r 3 -i 2, where 2 is its GPU ID. The DCGM documentation gives level 3 under 10 minutes on 4-GPU and under 35 minutes on 8-GPU Hopper systems; NVIDIA’s triage guide says the long test should take approximately 30 minutes. The PCIe test allows 80 replays per GPU by default. Which levels run on RTX PRO cards is covered in our acceptance testing guide.

Take, as an example, a server with eight RTX PRO 6000 Server Edition cards, four behind each PCIe switch, that logs 79 twice for one bus address. A third XID 79 at the new address under the same serial points at the card. XIDs on two cards behind one switch, with a bus fatal error in the BMC log, point at the switch path, and the server maker’s support comes first.

What a warranty case for XID 79 needs

Open the case once 79 returns after the checks above, when the catalogue’s investigatory action, contacting support, applies. Who handles it, the card’s supplier or the server maker, and what evidence it accepts follow that party’s warranty terms. NVIDIA’s debug guidelines list what a GPU issue report should contain and end with “Submit a ticket to your system vendor”. They also describe NVIDIA’s field diagnostic as “the authoritative and comprehensive NVIDIA tool for determining the health of GPUs”, which “is usually required before an RMA can be started”. NVIDIA leaves when and how to run it to the system vendor.

EVIDENCEHOW TO COLLECTASKED FOR BY
OS and driver detailsnvidia-bug-report.sh outputNVIDIA debug guidelines
Key log lines, full logkernel log with XID and AER linesNVIDIA debug guidelines
Debug steps takenthe troubleshooting order above, with resultsNVIDIA debug guidelines
DCGM diagnostics logdcgmi diag -r 3 with -j for JSONNVIDIA debug guidelines
BMC event logexport from the BMC, time of the XID notedour recommendation
Serial number and slotnvidia-smi -q and the label on the cardour recommendation
Field diagnosticrun only on the vendor’s instructionNVIDIA debug guidelines

NVIDIA GPU debug guidelines, GPU node triage, “Reporting a GPU Issue” and “Running Field Diagnostics”, updated 4 October 2026; rows marked «our recommendation» are ours.

For GPUs we supply, the manufacturer warranty is handled through us and support by our engineering partner’s service team. Send us the card, its serial number and the XID lines through the form below.

What we supply

We supply the cards this article discusses, the RTX PRO 6000 Server Edition and the H200 NVL with its NVLink bridges, and build AI servers to order around them, assembled and burn-in tested, with manufacturer warranty on every component. For the professional NVIDIA GPUs we supply, the manufacturer warranty is handled through us, DOA units are replaced, and support is handled by our engineering partner’s service team. When we build on your own chassis or with parts you already own, we check the platform, power and cooling first and say in advance if something will not work. Configuration and quote follow within one business day, on one EU contract and invoice.

FAQ

What does XID 79 mean?
XID 79 is the NVIDIA driver’s report that it tried to reach a GPU over its PCI Express connection and found it not accessible, logged as “GPU has fallen off the bus”. NVIDIA’s Xid catalogue gives a system restart as the immediate action and contacting support as the investigatory action. It names hardware failures on the PCIe link, failing GPU hardware and driver issues as the cause classes.
How do I fix XID 79 GPU has fallen off the bus?
Drain the node, collect nvidia-bug-report.sh output, the kernel log and the BMC event log, then restart the host, with a full power-off if the card stays missing. If the error returns, check the PCIe link and replays, the power cable and supplies, thermals, seating and risers, firmware and the driver in that order. A slot swap and DCGM diagnostics then show whether the card or the slot is at fault.
Why does nvidia-smi show no devices after XID 79?
nvidia-smi draws much of its function from NVIDIA’s management library, which returns NVML_ERROR_GPU_IS_LOST for a GPU that has fallen off the bus or otherwise become inaccessible. The card therefore fails or is missing in its output until the system is restarted, which is the action NVIDIA’s catalogue gives for XID 79. If it is still missing after a warm reboot, power the server off fully and check the card’s link and seating.
Can I reset the GPU instead of rebooting after XID 79?
NVIDIA’s catalogue gives a system restart, not a GPU reset, as the immediate action for XID 79. A reset with nvidia-smi -r needs root and no applications on the device, and NVIDIA states that it does not succeed in every case and is not recommended for production environments at this time. Plan for the host to go down and drain it first.
Is XID 79 a hardware failure?
NVIDIA writes that the event is often caused by hardware failures on the PCI Express link and may also come from failing GPU hardware or driver issues. A slot swap shows whether the error follows the card, stays with the slot or moves between GPUs, which points to the card, the slot path or the platform.
What do I need for an RMA after XID 79?
NVIDIA’s debug guidelines ask for the OS and driver details, the key log lines and the full log, the debug steps taken, nvidia-bug-report.sh output and DCGM diagnostics logs, submitted to the system vendor. NVIDIA’s field diagnostic is usually required before an RMA can start, and the system vendor says when and how to run it. Add the card’s serial number, its slot and the BMC event log, and follow the warranty terms of the supplier or manufacturer.

Send us the server model, the GPU cards and the slots they sit in, the XID lines with their PCI addresses and the power supplies fitted. We reply within one business day; for cards bought from us, the manufacturer warranty is handled through us.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna