BLOG · GUIDE ·

NVIDIA XID error codes in the XID Catalog: what 13, 31, 48, 74, 79, 94, 95 and 119 mean

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • An XID is an error report from the NVIDIA driver in the kernel log, a line with NVRM: Xid, the GPU’s PCI address and the code; NVIDIA’s XID Catalog, updated 9 September 2026, gives each code an immediate and an investigatory action
  • Application faults such as 13, 31 and 69 need a restart of the job; 64, 95, 119 and 120 have a GPU reset as the catalogue’s immediate action, and NVIDIA’s triage guide asks for a node reboot after 64; 79, a GPU that has fallen off the bus, calls for a system restart, a drained node and a report to the system vendor
  • XID 48 is an uncorrectable ECC error; if 63 or 64 follows, NVIDIA says to drain the node, wait for work to finish and reset the GPU, and 63 alone, a row remap, is marked IGNORE and takes effect at the next GPU reset
  • XID 94 and 95 appear only on GPUs with error containment; among the cards we supply, NVIDIA’s memory error document lists it for the H200 NVL, while the L40S, L4 and RTX Ada cards have row remapping only
  • Before a warranty case, collect nvidia-bug-report.sh output, the DCGM diagnostics log and the full kernel log; NVIDIA says its field diagnostic is usually required before an RMA can start

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What an NVIDIA XID error is

An XID is “an error report from the NVIDIA driver that is printed to the operating system’s kernel log or event log”, in NVIDIA’s words. On Linux each one is a line containing NVRM: Xid, the PCI address of the GPU, the error number and error-specific data. NVIDIA’s XID Catalog, updated on 9 September 2026, gives every code its trigger conditions, an immediate action and an investigatory action. The messages “can be indicative of a hardware problem, an NVIDIA software problem, or a user application problem”, and the code tells you where to look first.

The catalogue covers Ampere and newer GPUs, “including PCIe form-factor GPUs”, and is also published as the spreadsheet Xid-Catalog.xlsx. Its applicability columns name the A100, H100, B100 and GB200. The H200 NVL is a Hopper GPU like the H100. The L40S, L4 and the RTX cards are not named there, but as Ada and Blackwell GPUs they fall within the catalogue’s scope, except for codes that depend on a feature the card lacks, such as NVLink or error containment.

XID error codes list: meaning and NVIDIA’s action

The table lists the codes operators of PCIe GPU servers meet most, with NVIDIA’s resolution buckets.

XIDNVIDIA’S NAMELIKELY CAUSEIMMEDIATEINVESTIGATORY
13Graphics Engine Exceptionapplication fault, rarely hardwareRESTART_APPWORKFLOW_XID_13
31GPU memory page faultusually the application, sometimes driver or hardwareRESTART_APPWORKFLOW_XID_31
32Invalid or corrupted push buffer streamquality issues on PCIRESTART_APPCHECK_APP/CUDA
43GPU stopped processingapplication error; the GPU stays healthyIGNORECONTACT_SUPPORT
45Preemptive cleanup, due to previous errorsaborted application, often after another XIDWORKFLOW_XID_45alone: RESTART_FM; with others: follow them
48Double Bit ECC Erroruncorrectable memory errorWORKFLOW_XID_48WORKFLOW_XID_48
61PMU_BREAKPOINTmarked “Unused” in the catalogueCONTACT_SUPPORTnone listed
62Internal micro-controller haltmicro-controller on the GPURESET_GPUCONTACT_SUPPORT
63GPU memory remapping eventrow remapping after ECC errorsIGNOREIGNORE
64GPU memory remapping failurethe remap could not be recordedRESET_GPUCONTACT_SUPPORT
68NVDEC0 Exceptionvideo decoder engineRESTART_APPCONTACT_SUPPORT
69Graphics Engine class errorapplication or CUDARESTART_APPCHECK_APP/CUDA
74NVLINK Errorthe link, or the device at the remote endWORKFLOW_NVLINK_ERRCONTACT_SUPPORT
79GPU has fallen off the busPCIe link, failing GPU hardware, other driver issuesRESTART_BMCONTACT_SUPPORT
92High single-bit ECC error ratememory; no trigger text in the catalogueIGNORECONTACT_SUPPORT
94Contained memory errormemory, one application affectedRESTART_APPIGNORE (sympathetic)
95Uncontained memory errormemory, not contained; reset before restartRESET_GPUIGNORE (sympathetic)
119GSP RPC TimeoutGSP code, or no RPC reply in timeRESET_GPUINVESTIGATE_SW
120GSP Erroras 119RESET_GPUINVESTIGATE_SW
154GPU Recovery Action Changedsummary of the action another XID requiresXID_154none, informational only

NVIDIA XID Catalog, “Analyzing Xid Errors with the Xid Catalog”, updated 9 September 2026, and its PDF edition (release 615); the cause column is our short form of NVIDIA’s trigger conditions, and the action names are NVIDIA’s.

The PDF edition of the catalogue defines the buckets. RESTART_APP means “The application should be restarted”, RESTART_BM “Restart bare metal, system should be restarted” and IGNORE “No Action required”. RESET_GPU refers to NVIDIA’s GPU Debug Guidelines, updated 4 October 2026, for reset capabilities and limitations. For XID 79 the guide says “Drain and see Reporting a GPU Issue”, and for the WORKFLOW buckets it gives steps per code, which the next sections follow. It also says to report 61 and 62 and reset the GPU, while the catalogue marks 61 as unused.

How to find XID errors in the kernel log

NVIDIA writes that the messages “are logged in the kernel log buffer”, usually to the journal and from there to /var/log/messages or /var/log/syslog, and tells you to grep for “NVRM: Xid”. NVIDIA’s example line is NVRM: Xid (0000:03:00): 14, after a line that ties the PCI address to the GPU’s UUID. On a server with four or eight cards, work in this order.

  1. Search the current boot with journalctl -k | grep 'NVRM: Xid' or dmesg -T | grep 'NVRM: Xid'; earlier boots need a persistent journal (journalctl -k -b -1).
  2. Map each PCI address to a card with nvidia-smi and --query-gpu=index,pci.bus_id,serial plus --format=csv, so the case names the serial number, not a slot.
  3. Read the card’s state with nvidia-smi -q -d ECC,ROW_REMAPPER and nvidia-smi -q, which shows the GPU Recovery Action; NVIDIA lists None, Reset, Reboot, Drain P2P and Drain and Reset as its values.
  4. Check whether other GPUs logged codes at the same time, since 74 can point at the far end of a link.
  5. Run sudo nvidia-bug-report.sh before any reset or reboot; NVIDIA says to allow up to one hour, and to add --safe-mode --extra-system-data if it hangs.

Alert rules and DCGM fields are in our guide to GPU server monitoring with DCGM, and what the management controller reports out-of-band is in our article on out-of-band GPU monitoring through the BMC and Redfish.

Memory errors: XID 48, 63, 64, 92, 94 and 95

XID 48 is an uncorrectable, double-bit ECC error, and the catalogue states that “A GPU reset or node reboot is needed to clear this error.” The GPU Debug Guidelines set the order. If 63 or 64 follows, “Drain/cordon the node, wait for all work to complete, and reset GPU(s) reporting the XID.” If neither follows, the guide refers to NVIDIA’s field diagnostic to collect more debug information. XID 63 is the row remapper replacing a failing memory row with a spare, and NVIDIA says the remap “requires a GPU reset to take effect” and then stays for the life of the GPU. On its own, 63 has IGNORE in both buckets, and the remap waits for the next reset. XID 64 means the remap could not be recorded. The catalogue gives a GPU reset as the immediate action and a support case next, while the guide’s row for A100 says “The node should be rebooted immediately since there is a recording failure.” For XID 92, a high rate of single-bit errors, the guide also refers to the field diagnostic.

XIDs 94 and 95 are logged “when GPU drivers handle errors in GPUs that support error containment”. After 94 only the affected application restarts. After 95, “the affected GPU must be reset before applications can restart.” With MIG enabled, the guide says to drain the other GPU instances and reset; without MIG, the node “should be rebooted immediately”.

NVIDIA’s memory error document, updated 9 September 2026, lists the features by chip. The H200 NVL, a GH100, has error containment and row remapping, so all six codes can appear on it. The L40S, L4 and the RTX Ada cards are Ada AD10x chips with row remapping only, so 94 and 95 do not apply to them. The RTX PRO Blackwell cards are not in the table, although the nvidia-smi manual calls the row remapper “available on Ampere+”, so check what nvidia-smi -q -d ROW_REMAPPER reports on the delivered card. For a row remapping failure, NVIDIA’s RMA policy states that “the RMA criteria is met when the row-remapping failure flag is set and validated by the field diagnostic.”

NVLink, PCIe and GSP errors: XID 74, 79, 32, 119 and 120

XID 74 is an NVLink error. Of the cards we supply only the H200 NVL has NVLink, through bridges that span two or four cards. The catalogue says the error “may indicate a problem with the device at the remote end of the link”, so read the logs of every GPU on the same bridge; a GPU reset or node reboot “is needed to clear this error”. If it recurs, the guide decodes the first data word. Bits 4 or 5 point to a likely ECC or parity hardware issue, to be reported if seen more than twice on the same link; for bits 8, 9, 12, 16, 17, 21, 22, 24 and 28 it says to check the mechanical connections, and, outside 21 and 22, to re-seat if a field resolution is required and run diagnostics if the issue persists. From Ampere on, the nvidia-smi manual lets you reset one NVLink-connected GPU without its peers.

XID 79 means the driver tried to reach the GPU over PCI Express and could not. The catalogue says it is “often caused by hardware failures on the PCI Express link” and also names failing GPU hardware and other driver issues; system event logs and kernel PCI logs “may provide additional indications”. XID 32 points the same way, since its failures “primarily involve quality issues on PCI”. The checks in order are in our article on XID 79, the GPU that has fallen off the bus.

XIDs 119 and 120 come from the GSP, a processor on the GPU, with RESET_GPU and INVESTIGATE_SW as their buckets; the catalogue adds that “A GPU reset or node power cycle may be needed if the error persists.” For a software bucket, start with the driver: check your branch and its release notes, as our guide to NVIDIA driver branches and CUDA versions describes. The catalogue’s XID 154 column reads “CUDA 12.7; GPU driver R565” for 48, 62, 63, 64, 74, 79, 94, 95 and 120, and a 154 line names the new recovery action, such as Node Reboot Required.

Who acts on which XID

The resolution buckets also show who acts first.

NVIDIA BUCKETXIDSWHO ACTSFIRST STEP
RESTART_APP13, 31, 32, 68, 69, 94application ownerrestart the job; if 13 or 31 repeats, DCGM diagnostics rule out hardware; for 32, suspect PCIe
IGNORE43, 63, 92operator, records itlog it; 63 waits for the next reset, 92 gets the field diagnostic
RESET_GPU62, 64, 95, 119, 120operatordrain, reset the GPU or reboot, then check its health
RESTART_BM79operator, then warrantydrain the node, collect logs, restart, report
WORKFLOW45, 48, 74operatorfollow the guide’s steps for that code
CONTACT_SUPPORT43, 61, 62, 64, 68, 74, 79, 92vendor support; warranty once a fault is confirmedbug report, DCGM diagnostics log, field diagnostic

Buckets from NVIDIA’s XID Catalog; first steps from its GPU Debug Guidelines and the nvidia-smi manual; the roles are our example.

For a reset, the nvidia-smi manual says that nvidia-smi -r needs root and no applications on the GPU. It adds that a reset is not certain to work in all cases and “is not recommended for production environments at this time”, and says to power-cycle the node if a GPU is unhealthy afterwards. Before the card returns to service, run dcgmi diag -r 3, the long test. NVIDIA’s triage guide says it “should take approximately 30 minutes”, while the DCGM documentation gives under 10 minutes on 4-GPU Hopper systems.

Example, two inference servers with four H200 NVL each on one 4-way bridge. One GPU logs 48 followed by 63. The operator cordons that server, lets its requests move to the second one, waits for running work to end and resets the GPU; the level 3 diagnostic then decides whether the server goes back. After a 64, the guide asks for an immediate reboot, and a support case follows.

We assemble AI servers with these cards and load-test them before shipment, with a test report on request. Tell us how many servers and which cards you plan, and the configuration and quote follow within one business day.

What to send to the warranty provider

NVIDIA’s triage guide lists what a report to the vendor should contain; add when the error appeared and what changed, and compare with the results of your acceptance test and burn-in.

  1. Operating system, driver version and the server model.
  2. The XID lines and the key log messages, with the PCI address, serial number and slot of each card.
  3. A full listing of the log those lines come from, and the debug steps already taken.
  4. The nvidia-bug-report.log.gz file and the DCGM diagnostics log.
  5. Details of the application, if the error is tied to one.

NVIDIA calls its field diagnostic “the authoritative and comprehensive NVIDIA tool for determining the health of GPUs” and says it “is usually required before an RMA can be started”; the system vendor says when and how to run it.

For GPUs bought from us, the manufacturer warranty is handled through us. Send us the XID lines, the serial numbers and the bug report through the form below, and we reply within one business day.

What we supply

We supply the cards this reference covers: the H200 NVL with NVLink bridges, the L40S and L4, the RTX PRO 6000 Server Edition and the other RTX PRO Blackwell and RTX Ada GPUs, with manufacturer warranty, DOA units replaced and support from our engineering partner’s service team. We build AI servers to order with these cards, load-tested before shipment, with the operating system, drivers, CUDA and a container runtime installed on request, on one EU contract and invoice.

FAQ

What is an NVIDIA XID error?
An XID is an error report from the NVIDIA driver, printed to the kernel log as a line containing NVRM: Xid with the GPU’s PCI address and a code number. NVIDIA says it can point to a hardware problem, an NVIDIA software problem or a user application problem. The code tells you which, and NVIDIA’s XID Catalog gives the recommended action for each.
Where is the NVIDIA XID Catalog?
It is part of NVIDIA’s XID Errors documentation on docs.nvidia.com, under “Analyzing Xid Errors with the Xid Catalog”, and is also offered as the spreadsheet Xid-Catalog.xlsx. It covers Ampere and newer GPUs, including PCIe cards, and gives each code its trigger conditions, an immediate action and an investigatory action.
What does XID 79 mean?
XID 79 means the GPU has fallen off the bus: the driver tried to reach it over PCI Express and could not. NVIDIA’s catalogue names failures on the PCIe link, failing GPU hardware and other driver issues as causes, and its immediate action is a restart of the system. Its triage guide says to drain the node and report the issue to the system vendor.
What does XID 74 mean?
XID 74 is an NVLink error, a problem on a link to another GPU or at the device on the far end of it. A GPU reset or node reboot clears it. If it recurs, NVIDIA’s triage guide decodes the error data and, depending on the bits set, says to report it or to check the mechanical connections and re-seat the link, on an H200 NVL the bridge.
What do XID 48, 63 and 64 mean?
XID 48 is an uncorrectable double-bit ECC error, 63 a row remapping event and 64 a failure to record a remap. If 63 or 64 follows 48, NVIDIA says to drain the node, wait for work to finish and reset the GPU. A row remapping failure confirmed by NVIDIA’s field diagnostic meets its RMA criteria.
What does XID 119 mean?
XID 119 is a GSP RPC timeout, an error in code running on the GPU’s GSP core or a reply that did not arrive in time, and 120 is a related GSP error. NVIDIA’s catalogue gives a GPU reset as the immediate action and a software investigation as the next step. A GPU reset or node power cycle may be needed if the error persists.

Send us the XID lines from the kernel log, the card models and serial numbers, the server model and the driver version. We reply within one business day with the next step for each card and, for cards bought from us, how the manufacturer warranty case is handled through us.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna