NVIDIA XID error codes in the XID Catalog: what 13, 31, 48, 74, 79, 94, 95 and 119 mean
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- An XID is an error report from the NVIDIA driver in the kernel log, a line with NVRM: Xid, the GPU’s PCI address and the code; NVIDIA’s XID Catalog, updated 9 September 2026, gives each code an immediate and an investigatory action
- Application faults such as 13, 31 and 69 need a restart of the job; 64, 95, 119 and 120 have a GPU reset as the catalogue’s immediate action, and NVIDIA’s triage guide asks for a node reboot after 64; 79, a GPU that has fallen off the bus, calls for a system restart, a drained node and a report to the system vendor
- XID 48 is an uncorrectable ECC error; if 63 or 64 follows, NVIDIA says to drain the node, wait for work to finish and reset the GPU, and 63 alone, a row remap, is marked IGNORE and takes effect at the next GPU reset
- XID 94 and 95 appear only on GPUs with error containment; among the cards we supply, NVIDIA’s memory error document lists it for the H200 NVL, while the L40S, L4 and RTX Ada cards have row remapping only
- Before a warranty case, collect nvidia-bug-report.sh output, the DCGM diagnostics log and the full kernel log; NVIDIA says its field diagnostic is usually required before an RMA can start
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What an NVIDIA XID error is
An XID is “an error report from the NVIDIA driver that is printed to the operating system’s kernel log or event log”, in NVIDIA’s words. On Linux each one is a line containing NVRM: Xid, the PCI address of the GPU, the error number and error-specific data. NVIDIA’s XID Catalog, updated on 9 September 2026, gives every code its trigger conditions, an immediate action and an investigatory action. The messages “can be indicative of a hardware problem, an NVIDIA software problem, or a user application problem”, and the code tells you where to look first.
The catalogue covers Ampere and newer GPUs, “including PCIe form-factor GPUs”, and is also published as the spreadsheet Xid-Catalog.xlsx. Its applicability columns name the A100, H100, B100 and GB200. The H200 NVL is a Hopper GPU like the H100. The L40S, L4 and the RTX cards are not named there, but as Ada and Blackwell GPUs they fall within the catalogue’s scope, except for codes that depend on a feature the card lacks, such as NVLink or error containment.
XID error codes list: meaning and NVIDIA’s action
The table lists the codes operators of PCIe GPU servers meet most, with NVIDIA’s resolution buckets.
| XID | NVIDIA’S NAME | LIKELY CAUSE | IMMEDIATE | INVESTIGATORY |
|---|---|---|---|---|
| 13 | Graphics Engine Exception | application fault, rarely hardware | RESTART_APP | WORKFLOW_ |
| 31 | GPU memory page fault | usually the application, sometimes driver or hardware | RESTART_APP | WORKFLOW_ |
| 32 | Invalid or corrupted push buffer stream | quality issues on PCI | RESTART_APP | CHECK_ |
| 43 | GPU stopped processing | application error; the GPU stays healthy | IGNORE | CONTACT_ |
| 45 | Preemptive cleanup, due to previous errors | aborted application, often after another XID | WORKFLOW_ | alone: RESTART_FM; with others: follow them |
| 48 | Double Bit ECC Error | uncorrectable memory error | WORKFLOW_ | WORKFLOW_ |
| 61 | PMU_ | marked “Unused” in the catalogue | CONTACT_ | none listed |
| 62 | Internal micro-controller halt | micro-controller on the GPU | RESET_GPU | CONTACT_ |
| 63 | GPU memory remapping event | row remapping after ECC errors | IGNORE | IGNORE |
| 64 | GPU memory remapping failure | the remap could not be recorded | RESET_GPU | CONTACT_ |
| 68 | NVDEC0 Exception | video decoder engine | RESTART_APP | CONTACT_ |
| 69 | Graphics Engine class error | application or CUDA | RESTART_APP | CHECK_ |
| 74 | NVLINK Error | the link, or the device at the remote end | WORKFLOW_ | CONTACT_ |
| 79 | GPU has fallen off the bus | PCIe link, failing GPU hardware, other driver issues | RESTART_BM | CONTACT_ |
| 92 | High single-bit ECC error rate | memory; no trigger text in the catalogue | IGNORE | CONTACT_ |
| 94 | Contained memory error | memory, one application affected | RESTART_APP | IGNORE (sympathetic) |
| 95 | Uncontained memory error | memory, not contained; reset before restart | RESET_GPU | IGNORE (sympathetic) |
| 119 | GSP RPC Timeout | GSP code, or no RPC reply in time | RESET_GPU | INVESTIGATE_ |
| 120 | GSP Error | as 119 | RESET_GPU | INVESTIGATE_ |
| 154 | GPU Recovery Action Changed | summary of the action another XID requires | XID_154 | none, informational only |
NVIDIA XID Catalog, “Analyzing Xid Errors with the Xid Catalog”, updated 9 September 2026, and its PDF edition (release 615); the cause column is our short form of NVIDIA’s trigger conditions, and the action names are NVIDIA’s.
The PDF edition of the catalogue defines the buckets. RESTART_APP means “The application should be restarted”, RESTART_BM “Restart bare metal, system should be restarted” and IGNORE “No Action required”. RESET_GPU refers to NVIDIA’s GPU Debug Guidelines, updated 4 October 2026, for reset capabilities and limitations. For XID 79 the guide says “Drain and see Reporting a GPU Issue”, and for the WORKFLOW buckets it gives steps per code, which the next sections follow. It also says to report 61 and 62 and reset the GPU, while the catalogue marks 61 as unused.
How to find XID errors in the kernel log
NVIDIA writes that the messages “are logged in the kernel log buffer”, usually to the journal and from there to /var/log/messages or /var/log/syslog, and tells you to grep for “NVRM: Xid”. NVIDIA’s example line is NVRM: Xid (0000:03:00): 14, after a line that ties the PCI address to the GPU’s UUID. On a server with four or eight cards, work in this order.
- Search the current boot with
journalctl -k | grep 'NVRM: Xid'ordmesg -T | grep 'NVRM: Xid'; earlier boots need a persistent journal (journalctl -k -b -1). - Map each PCI address to a card with nvidia-smi and
--query-gpu=index,pci.bus_plusid,serial --format=csv, so the case names the serial number, not a slot. - Read the card’s state with
nvidia-smi -q -d ECC,ROW_andREMAPPER nvidia-smi -q, which shows the GPU Recovery Action; NVIDIA lists None, Reset, Reboot, Drain P2P and Drain and Reset as its values. - Check whether other GPUs logged codes at the same time, since 74 can point at the far end of a link.
- Run
sudo nvidia-bug-report.shbefore any reset or reboot; NVIDIA says to allow up to one hour, and to add--safe-mode --extra-system-dataif it hangs.
Alert rules and DCGM fields are in our guide to GPU server monitoring with DCGM, and what the management controller reports out-of-band is in our article on out-of-band GPU monitoring through the BMC and Redfish.
Memory errors: XID 48, 63, 64, 92, 94 and 95
XID 48 is an uncorrectable, double-bit ECC error, and the catalogue states that “A GPU reset or node reboot is needed to clear this error.” The GPU Debug Guidelines set the order. If 63 or 64 follows, “Drain/cordon the node, wait for all work to complete, and reset GPU(s) reporting the XID.” If neither follows, the guide refers to NVIDIA’s field diagnostic to collect more debug information. XID 63 is the row remapper replacing a failing memory row with a spare, and NVIDIA says the remap “requires a GPU reset to take effect” and then stays for the life of the GPU. On its own, 63 has IGNORE in both buckets, and the remap waits for the next reset. XID 64 means the remap could not be recorded. The catalogue gives a GPU reset as the immediate action and a support case next, while the guide’s row for A100 says “The node should be rebooted immediately since there is a recording failure.” For XID 92, a high rate of single-bit errors, the guide also refers to the field diagnostic.
XIDs 94 and 95 are logged “when GPU drivers handle errors in GPUs that support error containment”. After 94 only the affected application restarts. After 95, “the affected GPU must be reset before applications can restart.” With MIG enabled, the guide says to drain the other GPU instances and reset; without MIG, the node “should be rebooted immediately”.
NVIDIA’s memory error document, updated 9 September 2026, lists the features by chip. The H200 NVL, a GH100, has error containment and row remapping, so all six codes can appear on it. The L40S, L4 and the RTX Ada cards are Ada AD10x chips with row remapping only, so 94 and 95 do not apply to them. The RTX PRO Blackwell cards are not in the table, although the nvidia-smi manual calls the row remapper “available on Ampere+”, so check what nvidia-smi -q -d ROW_REMAPPER reports on the delivered card. For a row remapping failure, NVIDIA’s RMA policy states that “the RMA criteria is met when the row-remapping failure flag is set and validated by the field diagnostic.”
NVLink, PCIe and GSP errors: XID 74, 79, 32, 119 and 120
XID 74 is an NVLink error. Of the cards we supply only the H200 NVL has NVLink, through bridges that span two or four cards. The catalogue says the error “may indicate a problem with the device at the remote end of the link”, so read the logs of every GPU on the same bridge; a GPU reset or node reboot “is needed to clear this error”. If it recurs, the guide decodes the first data word. Bits 4 or 5 point to a likely ECC or parity hardware issue, to be reported if seen more than twice on the same link; for bits 8, 9, 12, 16, 17, 21, 22, 24 and 28 it says to check the mechanical connections, and, outside 21 and 22, to re-seat if a field resolution is required and run diagnostics if the issue persists. From Ampere on, the nvidia-smi manual lets you reset one NVLink-connected GPU without its peers.
XID 79 means the driver tried to reach the GPU over PCI Express and could not. The catalogue says it is “often caused by hardware failures on the PCI Express link” and also names failing GPU hardware and other driver issues; system event logs and kernel PCI logs “may provide additional indications”. XID 32 points the same way, since its failures “primarily involve quality issues on PCI”. The checks in order are in our article on XID 79, the GPU that has fallen off the bus.
XIDs 119 and 120 come from the GSP, a processor on the GPU, with RESET_GPU and INVESTIGATE_SW as their buckets; the catalogue adds that “A GPU reset or node power cycle may be needed if the error persists.” For a software bucket, start with the driver: check your branch and its release notes, as our guide to NVIDIA driver branches and CUDA versions describes. The catalogue’s XID 154 column reads “CUDA 12.7; GPU driver R565” for 48, 62, 63, 64, 74, 79, 94, 95 and 120, and a 154 line names the new recovery action, such as Node Reboot Required.
Who acts on which XID
The resolution buckets also show who acts first.
| NVIDIA BUCKET | XIDS | WHO ACTS | FIRST STEP |
|---|---|---|---|
| RESTART_APP | 13, 31, 32, 68, 69, 94 | application owner | restart the job; if 13 or 31 repeats, DCGM diagnostics rule out hardware; for 32, suspect PCIe |
| IGNORE | 43, 63, 92 | operator, records it | log it; 63 waits for the next reset, 92 gets the field diagnostic |
| RESET_GPU | 62, 64, 95, 119, 120 | operator | drain, reset the GPU or reboot, then check its health |
| RESTART_BM | 79 | operator, then warranty | drain the node, collect logs, restart, report |
| WORKFLOW | 45, 48, 74 | operator | follow the guide’s steps for that code |
| CONTACT_ | 43, 61, 62, 64, 68, 74, 79, 92 | vendor support; warranty once a fault is confirmed | bug report, DCGM diagnostics log, field diagnostic |
Buckets from NVIDIA’s XID Catalog; first steps from its GPU Debug Guidelines and the nvidia-smi manual; the roles are our example.
For a reset, the nvidia-smi manual says that nvidia-smi -r needs root and no applications on the GPU. It adds that a reset is not certain to work in all cases and “is not recommended for production environments at this time”, and says to power-cycle the node if a GPU is unhealthy afterwards. Before the card returns to service, run dcgmi diag -r 3, the long test. NVIDIA’s triage guide says it “should take approximately 30 minutes”, while the DCGM documentation gives under 10 minutes on 4-GPU Hopper systems.
Example, two inference servers with four H200 NVL each on one 4-way bridge. One GPU logs 48 followed by 63. The operator cordons that server, lets its requests move to the second one, waits for running work to end and resets the GPU; the level 3 diagnostic then decides whether the server goes back. After a 64, the guide asks for an immediate reboot, and a support case follows.
We assemble AI servers with these cards and load-test them before shipment, with a test report on request. Tell us how many servers and which cards you plan, and the configuration and quote follow within one business day.
What to send to the warranty provider
NVIDIA’s triage guide lists what a report to the vendor should contain; add when the error appeared and what changed, and compare with the results of your acceptance test and burn-in.
- Operating system, driver version and the server model.
- The XID lines and the key log messages, with the PCI address, serial number and slot of each card.
- A full listing of the log those lines come from, and the debug steps already taken.
- The
nvidia-bug-report.log.gzfile and the DCGM diagnostics log. - Details of the application, if the error is tied to one.
NVIDIA calls its field diagnostic “the authoritative and comprehensive NVIDIA tool for determining the health of GPUs” and says it “is usually required before an RMA can be started”; the system vendor says when and how to run it.
For GPUs bought from us, the manufacturer warranty is handled through us. Send us the XID lines, the serial numbers and the bug report through the form below, and we reply within one business day.
What we supply
We supply the cards this reference covers: the H200 NVL with NVLink bridges, the L40S and L4, the RTX PRO 6000 Server Edition and the other RTX PRO Blackwell and RTX Ada GPUs, with manufacturer warranty, DOA units replaced and support from our engineering partner’s service team. We build AI servers to order with these cards, load-tested before shipment, with the operating system, drivers, CUDA and a container runtime installed on request, on one EU contract and invoice.
FAQ
What is an NVIDIA XID error?
Where is the NVIDIA XID Catalog?
What does XID 79 mean?
What does XID 74 mean?
What do XID 48, 63 and 64 mean?
What does XID 119 mean?
Send us the XID lines from the kernel log, the card models and serial numbers, the server model and the driver version. We reply within one business day with the next step for each card and, for cards bought from us, how the manufacturer warranty case is handled through us.
Talk to an expertWe reply within one business day