BLOG · GUIDE ·

GPU out-of-band monitoring: what the BMC and Redfish show, and what only DCGM sees in-band

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Out-of-band GPU monitoring reads the cards through the server’s BMC on its own management port, without the operating system; NVIDIA’s RTX PRO 6000 Server Edition brief adds that its SMBPBI power limit channel needs the driver loaded for full functionality, and Xid errors stay in the kernel log, in-band
  • NVIDIA’s H200 NVL and RTX PRO 6000 Server Edition briefs list SMBPBI, the SMBus Post-Box Interface, as supported and say a power cap set through it stays in force across driver loads and boots, while one set with nvidia-smi must be set again after each driver load
  • Redfish 2026.2 describes a GPU as a Processor of type GPU, with EnvironmentMetrics for temperature, power and power limit and ProcessorMetrics for bandwidth, error counts and throttle durations; each BMC decides which of them it fills
  • Dell’s iDRAC telemetry lists 40 GPU metrics, SM activity and NVLink error bits among them, plus 16 memory error counters, under an iDRAC “Datacenter” licence; Lenovo XCC2 documents a GPU power sensor and power limit, Supermicro a GPU inventory tab
  • For 2 to 4 GPU servers, send BMC events and DCGM metrics into one monitoring system, and note in a maintenance window which GPU values the BMC still reports with the NVIDIA driver unloaded

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

GPU out-of-band monitoring: what the BMC sees and what it does not

Out-of-band GPU monitoring reads the GPUs through the server’s baseboard management controller (BMC), on its own management port, instead of through the operating system. On data-centre cards the BMC reaches the GPU over the SMBus, through NVIDIA’s SMBus Post-Box Interface (SMBPBI) where the server maker implements it, and the makers document temperature, board power and health data in Redfish and their web consoles. The path is not fully independent of the driver, though: NVIDIA’s RTX PRO 6000 Server Edition brief says its SMBPBI channel for the power limit “also requires that the NVIDIA driver is loaded for full functionality”. The detail that explains a slow or failed job, Xid errors above all, comes in-band from the driver, through nvidia-smi and NVIDIA’s Data Center GPU Manager (DCGM), and stops when the host or the driver stops.

How much the BMC shows beyond temperature and power depends on the server maker. Dell’s iDRAC telemetry lists 40 GPU metrics, including SM activity and NVLink error flags, while other makers document little beyond inventory and power for PCIe GPUs. A site with 2 to 4 GPU servers therefore needs both paths in one monitoring system: the BMC for the readings that must survive a hung host, and DCGM for the counters that only the driver has. Documents cited are as of October 2026.

How the BMC reads a PCIe GPU: SMBus and SMBPBI

A data-centre GPU has its own management interface on the SMBus. NVIDIA’s H200 NVL product brief, PB-12128-001_v01 of 11 April 2025, lists the SMBus addresses 0x9E for writes and 0x9F for reads, “SMBus direct access Supported” and “SMBPBI (SMBus Post-Box Interface) Supported”. Of the NVIDIA briefs we read, only the RTX PRO 6000 Server Edition brief lists an SMBPBI command, the request to set the total GPU power limit. None lists the telemetry values SMBPBI returns, so the BMC firmware of each server maker decides which values it reads and how it presents them.

The power cap shows the difference between the two paths. In the H200 NVL brief, a cap set out-of-band through SMBPBI “remains in force across driver loads and system boots”, while one set in-band with nvidia-smi “must be reestablished after each new driver load”. NVIDIA rates the H200 NVL for a 200 W minimum and a 600 W maximum, the default. For four H200 NVL on a limited rack feed, a cap set through the BMC survives driver updates, while an nvidia-smi cap needs a service that sets it at every boot.

The L4 brief, PB-11316-001_v01 of 9 March 2023, lists SMBus direct access and SMBPBI as supported. The RTX PRO 6000 Server Edition brief, SP-12355-001_v02 of 27 June 2025, lists SMBPBI as supported and says the card “supports GPU out-of-band telemetry and firmware updates through SMBus and USB 2.0”. It repeats the power cap wording of the H200 NVL brief, with a 300 W minimum. For the L40S, check its own brief, and for every card the server maker’s list of supported GPUs, before relying on BMC readings.

GPU data in Redfish: Processor, EnvironmentMetrics and ProcessorMetrics

Redfish is the DMTF’s REST interface for server management, and its schema bundle 2026.2 was published on 14 September 2026. A GPU appears as a Processor resource whose ProcessorType is GPU; the schema also offers Accelerator. The Processor carries Status, “the status and health of the resource”, Throttled and ThrottleCauses, and a memory summary with ECCModeEnabled. It links to two metric resources.

EnvironmentMetrics, “the environmental metrics of a device”, holds TemperatureCelsius, PowerWatts, PowerLimitWatts and EnergykWh. ProcessorMetrics, “usage and health statistics for a processor”, holds BandwidthPercent, OperatingSpeedMHz, correctable and uncorrectable error counts, PCIe errors, and PowerLimitThrottleDuration and ThermalLimitThrottleDuration, the time spent throttling since reset. The BMC firmware decides which of these properties it fills, so read the Processor resources of one delivered server before writing alert rules.

METRICOUT-OF-BAND, BMCIN-BAND, DRIVER
GPU temperatureRedfish TemperatureCelsius; Dell PrimaryTemperature and MemoryTemperaturenvidia-smi, DCGM temperature fields
Board powerRedfish PowerWatts; Dell PowerConsumption; Lenovo GPU{N}_Power sensornvidia-smi, DCGM power field
Power capSMBPBI on the H200 NVL and RTX PRO 6000 Server Edition, kept across driver loads; Lenovo GPU{N}_PowerLimitnvidia-smi, set again after each driver load
ThrottlingRedfish ThrottleCauses and throttle durations; Dell clock event reason bitsnvidia-smi clock event reasons, DCGM violation counters
Utilisation, SM activityRedfish BandwidthPercent (bandwidth use); Dell GPUUsage and GPUSMActivityDCGM profiling fields
Memory errorsDell GPU Statistics counters, sensed every 600 snvidia-smi ECC and row remapper, DCGM ECC fields
NVLinkDell link status and runtime error bitsnvidia-smi nvlink, DCGM NVLink counters
Xid errorsnone in the BMC documents we readkernel log, DCGM, dcgm-exporter
Inventory, firmwareSupermicro GPU tab with model, serial, part number and firmwarenvidia-smi -q with VBIOS version

Redfish schemas Processor v1_24_0, EnvironmentMetrics v1_7_0 and ProcessorMetrics v1_7_0 (DMTF); Dell iDRAC Telemetry Reference Guide; Lenovo XCC2 REST API; Supermicro BMC manual X14/H14 rev. 1.1; NVIDIA H200 NVL and RTX PRO 6000 Server Edition product briefs; NVIDIA DCGM 4.6 documentation and Xid documentation, all read on 10 October 2026.

What Dell, HPE, Lenovo and Supermicro document for GPUs

The four makers document very different depths for PCIe GPUs. The table lists only what each document states; a property missing from it may still exist on your model and firmware.

SERVER MAKER, BMCGPU DATA DOCUMENTEDCONDITIONS STATED
Dell iDRAC9 and iDRAC10Telemetry reports GPU Metrics (40 metrics, mostly every 5 s) and GPU Statistics (16 error counters, every 600 s)iDRAC “Datacenter” licence; 14th generation or newer; iDRAC 4.0 or higher per Dell’s tools README
HPE iLO 6Processors/{id}/EnvironmentMetrics and ProcessorMetrics added in v1.61; PowerWatts in v1.68; power limit by PATCH in v1.75changelog does not say which GPU models fill them
Lenovo XCC2Controls/GPU{N}_PowerLimit with the sensor GPU{N}_Power, linked to the GPU’s Processor resourcesetting the limit needs the XCC2 Platinum licence; on AMD-based systems NVIDIA GPUs only
Supermicro X14/H14 BMCComponent Information, GPU tab: vendor and model, serial number, part number, firmware versionfeature table lists “GPU monitoring (NVIDIA GPUs)”

Dell iDRAC Telemetry Reference Guide (GPU Metrics, GPU Statistics) and Dell’s iDRAC-Telemetry-Reference-Tools README; HPE iLO 6 Redfish changelog up to v1.79; Lenovo XCC2 REST API, GET GPU PowerLimit properties; Supermicro BMC manual X14/H14, revision 1.1, 22 December 2025. Read on 10 October 2026.

Dell’s reference lists far more GPU values than the other three documents. Its GPU Metrics report includes GPUResetRecommendedState, described as “A flag that indicates if a GPU reset is recommended”, PowerBrakeState, ThermalAlertState, PCIe correctable error counts and per-link NVLink error bits. GPU Statistics counts single-bit and double-bit errors per memory region and the pages retired because of them. Telemetry, streamed or pulled, needs the iDRAC “Datacenter” licence, so put it in the order. Dell’s guide does not say how iDRAC collects these values.

Lenovo’s XCC2 documents GPU power as a Redfish control: GPU{N}_PowerLimit reads the sensor GPU{N}_Power and takes a set point in watts, hidden without the Platinum licence. HPE’s iLO 6 changelog adds processor EnvironmentMetrics and ProcessorMetrics in v1.61. Its GPU entries name an NVLink state between CPU and GPU and the values GPU1 and GPU2 in two power schemas, but no GPU models, so for a PCIe card ask HPE which properties your model fills. Supermicro’s X14/H14 manual documents a GPU inventory tab with firmware versions, useful when an update reaches only some cards.

We build AI servers to order with out-of-band management over IPMI and the BMC. Tell us which server maker you prefer and which monitoring system should read the BMC.

Xid errors, ECC detail and NVLink: the in-band path

NVIDIA defines the Xid message as “an error report from the NVIDIA driver that is printed to the operating system’s kernel log or event log”, in its Xid documentation updated on 9 September 2026. An Xid exists only where the driver runs, so it reaches the monitoring system through the kernel log or DCGM. DCGM itself works through NVML, the driver’s management library: its 4.6 documentation describes how dcgmDetachDriver “detaches NVML from DCGM”. When the driver stops, the in-band path stops with it.

In-band, nvidia-smi and DCGM give the full detail behind a GPU reset, a drained node or a warranty case: Xid codes, volatile and aggregate ECC counts, row remapping, clock event reasons per sample and NVLink replay and CRC counts. Dell’s telemetry reports part of this out-of-band, with single-bit and double-bit error counts every 600 seconds, clock event reason bits and NVLink error bits, but none of the BMC documents we read lists row remapping or Xid codes. Our guide to DCGM metrics and XID errors covers the fields and the dcgm-exporter defaults, and the NVIDIA Xid error codes reference lists the codes with NVIDIA’s actions. For a card the driver can no longer reach, see our article on Xid 79, a GPU fallen off the bus, and keep the BMC’s readings for that slot as a second record for the warranty case.

The two paths also differ in what keeps running during maintenance. A driver update unloads the module, and dcgm-exporter reports nothing until it is back. The BMC keeps running, but none of the documents above states which GPU values it still reads without the driver, so check that during the first driver update. Inference latency, queue length and KV cache use sit one layer higher again, in the serving engine, as our LLM serving monitoring guide explains.

Forwarding BMC alerts to the monitoring system

Redfish offers two ways to get data out of the BMC without polling every resource. The EventService “contains properties for managing event subscriptions and generates the events sent to subscribers”; each subscription is an event destination, and the service can also offer a Server-Sent Events stream and SMTP delivery. The TelemetryService “is used for collecting and reporting metric data within the Redfish service”, with metric report definitions and metric reports; Dell’s GPU Metrics and GPU Statistics are such reports. Older interfaces remain, and Supermicro’s X14/H14 manual lists alerts, SNMP, syslog and SMTP under its notification settings.

For a few servers, one collector per site that subscribes to each BMC’s events and polls the GPU temperature and power resources is enough. Label every series with the server, the slot and the card’s serial number, so that a BMC alert and a DCGM alert for the same card meet in one incident. The BMC should sit on an isolated management network that only the collector and a jump host reach; our guide to securing a GPU server covers that network and BMC firmware.

Setting up both paths on 2 to 4 GPU servers

As an example, take three servers with eight RTX PRO 6000 Server Edition cards each and one with four H200 NVL, 28 GPUs in all. If these were Dell models with the “Datacenter” licence, the GPU Metrics report alone would deliver 40 values per GPU, 1,120 for the site, most of them every 5 seconds, by our arithmetic. This order sets up both paths and tests what each one reports.

  1. Connect each BMC to the management network, update its firmware to the maker’s current release and record the GPU inventory it shows against nvidia-smi -q.
  2. Read the Processor resources of type GPU on one server and note which EnvironmentMetrics and ProcessorMetrics properties it fills.
  3. Subscribe the collector to each BMC’s Redfish events, or configure SNMP or syslog where Redfish events are not available.
  4. Poll GPU temperature and power out-of-band, and set the thresholds from each card’s slowdown temperature and power limit.
  5. Install dcgm-exporter in-band, ship the kernel log for Xid lines, and add the ECC, row remapping and NVLink fields your cards support.
  6. In a maintenance window, unload the NVIDIA driver on one server, note which GPU values the BMC still reports and confirm that a test temperature alert from it arrives; then load the driver and confirm that an in-band test alert arrives too.
  7. Write down which path owns which alert, so that whoever takes the alert knows where to look first.

Our AI servers are assembled and burn-in tested before shipment, with a test report on request. Send us the cards and server count you plan through the form below, and we reply within one business day with a configuration and quote.

What we supply

We build AI servers to order with out-of-band management over IPMI and the BMC, assembled and burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice. The cards in this article come from our GPU range: the H200 NVL with NVLink bridges, the RTX PRO 6000 Server Edition, the L40S and the L4. The operating system, drivers, CUDA and a container runtime are installed on request, and support for the cards is handled by our engineering partner’s service team. If your monitoring relies on a BMC feature that needs a licence, such as Dell’s iDRAC “Datacenter” licence for telemetry, name it in your request.

FAQ

What is out-of-band GPU monitoring?
It reads the GPUs through the server’s baseboard management controller on its own management port, instead of through the operating system. On data-centre cards the BMC reaches the GPU over the SMBus, through NVIDIA’s SMBus Post-Box Interface where the server maker implements it, and exposes temperature, board power and health through Redfish and its web console. Xid errors and row remapping stay in-band, in the driver and DCGM, while some BMCs, Dell’s iDRAC among them, also report ECC counts, clock event reasons and NVLink error flags.
Can the BMC monitor GPU temperature?
Yes, on servers whose BMC firmware supports the installed GPU model; NVIDIA’s RTX PRO 6000 Server Edition brief says that card reports its temperature values over in-band and out-of-band management paths. Redfish describes GPU temperature as TemperatureCelsius in the EnvironmentMetrics of the GPU’s Processor resource, and Dell’s iDRAC telemetry lists primary, memory and board temperatures for GPUs. Check the server maker’s supported GPU list and read the values on one delivered server before setting thresholds.
What GPU telemetry does Redfish provide?
Redfish 2026.2 models a GPU as a Processor of type GPU, with status, throttle state and causes, and a memory summary with ECC mode. Its EnvironmentMetrics hold temperature, power, power limit and energy, and its ProcessorMetrics hold bandwidth use, operating speed, error counts and throttle durations. The schema defines these properties, while each BMC decides which of them it fills.
Does iDRAC monitor NVIDIA GPUs?
Dell’s iDRAC Telemetry Reference Guide, which covers iDRAC9 and iDRAC10, lists a GPU Metrics report with 40 metrics, including temperatures, board power, power caps, clock event reasons, SM activity, PCIe and NVLink errors, and a GPU Statistics report with 16 memory error counters. Telemetry requires an iDRAC “Datacenter” licence and a 14th-generation or newer server.
What is NVIDIA SMBPBI?
SMBPBI is NVIDIA’s SMBus Post-Box Interface on data-centre GPUs, through which a server’s BMC communicates with the card over the SMBus, without going through the operating system. NVIDIA’s H200 NVL, L4 and RTX PRO 6000 Server Edition product briefs list it as supported, and the H200 NVL and RTX PRO 6000 Server Edition briefs state that a power cap set through SMBPBI remains in force across driver loads and system boots. The briefs do not list the telemetry values it returns, so the values shown depend on the server maker’s BMC firmware.
Are NVIDIA Xid errors visible out-of-band?
Not in the BMC documents of Dell, HPE, Lenovo and Supermicro that we read in October 2026. NVIDIA defines an Xid as an error report from the driver printed to the operating system’s kernel log or event log, so it is collected in-band from that log or through DCGM. Dell’s iDRAC does offer a reset-recommended flag per GPU out-of-band, which can complement the Xid alert.

Send us the number of GPU servers and cards you plan, the server maker you prefer and the monitoring system that should receive BMC events and DCGM metrics. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna