GPU out-of-band monitoring: what the BMC and Redfish show, and what only DCGM sees in-band
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Out-of-band GPU monitoring reads the cards through the server’s BMC on its own management port, without the operating system; NVIDIA’s RTX PRO 6000 Server Edition brief adds that its SMBPBI power limit channel needs the driver loaded for full functionality, and Xid errors stay in the kernel log, in-band
- NVIDIA’s H200 NVL and RTX PRO 6000 Server Edition briefs list SMBPBI, the SMBus Post-Box Interface, as supported and say a power cap set through it stays in force across driver loads and boots, while one set with nvidia-smi must be set again after each driver load
- Redfish 2026.2 describes a GPU as a Processor of type GPU, with EnvironmentMetrics for temperature, power and power limit and ProcessorMetrics for bandwidth, error counts and throttle durations; each BMC decides which of them it fills
- Dell’s iDRAC telemetry lists 40 GPU metrics, SM activity and NVLink error bits among them, plus 16 memory error counters, under an iDRAC “Datacenter” licence; Lenovo XCC2 documents a GPU power sensor and power limit, Supermicro a GPU inventory tab
- For 2 to 4 GPU servers, send BMC events and DCGM metrics into one monitoring system, and note in a maintenance window which GPU values the BMC still reports with the NVIDIA driver unloaded
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
GPU out-of-band monitoring: what the BMC sees and what it does not
Out-of-band GPU monitoring reads the GPUs through the server’s baseboard management controller (BMC), on its own management port, instead of through the operating system. On data-centre cards the BMC reaches the GPU over the SMBus, through NVIDIA’s SMBus Post-Box Interface (SMBPBI) where the server maker implements it, and the makers document temperature, board power and health data in Redfish and their web consoles. The path is not fully independent of the driver, though: NVIDIA’s RTX PRO 6000 Server Edition brief says its SMBPBI channel for the power limit “also requires that the NVIDIA driver is loaded for full functionality”. The detail that explains a slow or failed job, Xid errors above all, comes in-band from the driver, through nvidia-smi and NVIDIA’s Data Center GPU Manager (DCGM), and stops when the host or the driver stops.
How much the BMC shows beyond temperature and power depends on the server maker. Dell’s iDRAC telemetry lists 40 GPU metrics, including SM activity and NVLink error flags, while other makers document little beyond inventory and power for PCIe GPUs. A site with 2 to 4 GPU servers therefore needs both paths in one monitoring system: the BMC for the readings that must survive a hung host, and DCGM for the counters that only the driver has. Documents cited are as of October 2026.
How the BMC reads a PCIe GPU: SMBus and SMBPBI
A data-centre GPU has its own management interface on the SMBus. NVIDIA’s H200 NVL product brief, PB-12128-001_v01 of 11 April 2025, lists the SMBus addresses 0x9E for writes and 0x9F for reads, “SMBus direct access Supported” and “SMBPBI (SMBus Post-Box Interface) Supported”. Of the NVIDIA briefs we read, only the RTX PRO 6000 Server Edition brief lists an SMBPBI command, the request to set the total GPU power limit. None lists the telemetry values SMBPBI returns, so the BMC firmware of each server maker decides which values it reads and how it presents them.
The power cap shows the difference between the two paths. In the H200 NVL brief, a cap set out-of-band through SMBPBI “remains in force across driver loads and system boots”, while one set in-band with nvidia-smi “must be reestablished after each new driver load”. NVIDIA rates the H200 NVL for a 200 W minimum and a 600 W maximum, the default. For four H200 NVL on a limited rack feed, a cap set through the BMC survives driver updates, while an nvidia-smi cap needs a service that sets it at every boot.
The L4 brief, PB-11316-001_v01 of 9 March 2023, lists SMBus direct access and SMBPBI as supported. The RTX PRO 6000 Server Edition brief, SP-12355-001_v02 of 27 June 2025, lists SMBPBI as supported and says the card “supports GPU out-of-band telemetry and firmware updates through SMBus and USB 2.0”. It repeats the power cap wording of the H200 NVL brief, with a 300 W minimum. For the L40S, check its own brief, and for every card the server maker’s list of supported GPUs, before relying on BMC readings.
GPU data in Redfish: Processor, EnvironmentMetrics and ProcessorMetrics
Redfish is the DMTF’s REST interface for server management, and its schema bundle 2026.2 was published on 14 September 2026. A GPU appears as a Processor resource whose ProcessorType is GPU; the schema also offers Accelerator. The Processor carries Status, “the status and health of the resource”, Throttled and ThrottleCauses, and a memory summary with ECCModeEnabled. It links to two metric resources.
EnvironmentMetrics, “the environmental metrics of a device”, holds TemperatureCelsius, PowerWatts, PowerLimitWatts and EnergykWh. ProcessorMetrics, “usage and health statistics for a processor”, holds BandwidthPercent, OperatingSpeedMHz, correctable and uncorrectable error counts, PCIe errors, and PowerLimitThrottleDuration and ThermalLimitThrottleDuration, the time spent throttling since reset. The BMC firmware decides which of these properties it fills, so read the Processor resources of one delivered server before writing alert rules.
| METRIC | OUT-OF-BAND, BMC | IN-BAND, DRIVER |
|---|---|---|
| GPU temperature | Redfish Temperature | nvidia-smi, DCGM temperature fields |
| Board power | Redfish PowerWatts; Dell Power | nvidia-smi, DCGM power field |
| Power cap | SMBPBI on the H200 NVL and RTX PRO 6000 Server Edition, kept across driver loads; Lenovo GPU{N}_ | nvidia-smi, set again after each driver load |
| Throttling | Redfish Throttle | nvidia-smi clock event reasons, DCGM violation counters |
| Utilisation, SM activity | Redfish Bandwidth | DCGM profiling fields |
| Memory errors | Dell GPU Statistics counters, sensed every 600 s | nvidia-smi ECC and row remapper, DCGM ECC fields |
| NVLink | Dell link status and runtime error bits | nvidia-smi nvlink, DCGM NVLink counters |
| Xid errors | none in the BMC documents we read | kernel log, DCGM, dcgm-exporter |
| Inventory, firmware | Supermicro GPU tab with model, serial, part number and firmware | nvidia-smi -q with VBIOS version |
Redfish schemas Processor v1_24_0, EnvironmentMetrics v1_7_0 and ProcessorMetrics v1_7_0 (DMTF); Dell iDRAC Telemetry Reference Guide; Lenovo XCC2 REST API; Supermicro BMC manual X14/H14 rev. 1.1; NVIDIA H200 NVL and RTX PRO 6000 Server Edition product briefs; NVIDIA DCGM 4.6 documentation and Xid documentation, all read on 10 October 2026.
What Dell, HPE, Lenovo and Supermicro document for GPUs
The four makers document very different depths for PCIe GPUs. The table lists only what each document states; a property missing from it may still exist on your model and firmware.
| SERVER MAKER, BMC | GPU DATA DOCUMENTED | CONDITIONS STATED |
|---|---|---|
| Dell iDRAC9 and iDRAC10 | Telemetry reports GPU Metrics (40 metrics, mostly every 5 s) and GPU Statistics (16 error counters, every 600 s) | iDRAC “Datacenter” licence; 14th generation or newer; iDRAC 4.0 or higher per Dell’s tools README |
| HPE iLO 6 | Processors/ | changelog does not say which GPU models fill them |
| Lenovo XCC2 | Controls/ | setting the limit needs the XCC2 Platinum licence; on AMD-based systems NVIDIA GPUs only |
| Supermicro X14/H14 BMC | Component Information, GPU tab: vendor and model, serial number, part number, firmware version | feature table lists “GPU monitoring (NVIDIA GPUs)” |
Dell iDRAC Telemetry Reference Guide (GPU Metrics, GPU Statistics) and Dell’s iDRAC-Telemetry-Reference-Tools README; HPE iLO 6 Redfish changelog up to v1.79; Lenovo XCC2 REST API, GET GPU PowerLimit properties; Supermicro BMC manual X14/H14, revision 1.1, 22 December 2025. Read on 10 October 2026.
Dell’s reference lists far more GPU values than the other three documents. Its GPU Metrics report includes GPUResetRecommendedState, described as “A flag that indicates if a GPU reset is recommended”, PowerBrakeState, ThermalAlertState, PCIe correctable error counts and per-link NVLink error bits. GPU Statistics counts single-bit and double-bit errors per memory region and the pages retired because of them. Telemetry, streamed or pulled, needs the iDRAC “Datacenter” licence, so put it in the order. Dell’s guide does not say how iDRAC collects these values.
Lenovo’s XCC2 documents GPU power as a Redfish control: GPU{N}_PowerLimit reads the sensor GPU{N}_Power and takes a set point in watts, hidden without the Platinum licence. HPE’s iLO 6 changelog adds processor EnvironmentMetrics and ProcessorMetrics in v1.61. Its GPU entries name an NVLink state between CPU and GPU and the values GPU1 and GPU2 in two power schemas, but no GPU models, so for a PCIe card ask HPE which properties your model fills. Supermicro’s X14/H14 manual documents a GPU inventory tab with firmware versions, useful when an update reaches only some cards.
We build AI servers to order with out-of-band management over IPMI and the BMC. Tell us which server maker you prefer and which monitoring system should read the BMC.
Xid errors, ECC detail and NVLink: the in-band path
NVIDIA defines the Xid message as “an error report from the NVIDIA driver that is printed to the operating system’s kernel log or event log”, in its Xid documentation updated on 9 September 2026. An Xid exists only where the driver runs, so it reaches the monitoring system through the kernel log or DCGM. DCGM itself works through NVML, the driver’s management library: its 4.6 documentation describes how dcgmDetachDriver “detaches NVML from DCGM”. When the driver stops, the in-band path stops with it.
In-band, nvidia-smi and DCGM give the full detail behind a GPU reset, a drained node or a warranty case: Xid codes, volatile and aggregate ECC counts, row remapping, clock event reasons per sample and NVLink replay and CRC counts. Dell’s telemetry reports part of this out-of-band, with single-bit and double-bit error counts every 600 seconds, clock event reason bits and NVLink error bits, but none of the BMC documents we read lists row remapping or Xid codes. Our guide to DCGM metrics and XID errors covers the fields and the dcgm-exporter defaults, and the NVIDIA Xid error codes reference lists the codes with NVIDIA’s actions. For a card the driver can no longer reach, see our article on Xid 79, a GPU fallen off the bus, and keep the BMC’s readings for that slot as a second record for the warranty case.
The two paths also differ in what keeps running during maintenance. A driver update unloads the module, and dcgm-exporter reports nothing until it is back. The BMC keeps running, but none of the documents above states which GPU values it still reads without the driver, so check that during the first driver update. Inference latency, queue length and KV cache use sit one layer higher again, in the serving engine, as our LLM serving monitoring guide explains.
Forwarding BMC alerts to the monitoring system
Redfish offers two ways to get data out of the BMC without polling every resource. The EventService “contains properties for managing event subscriptions and generates the events sent to subscribers”; each subscription is an event destination, and the service can also offer a Server-Sent Events stream and SMTP delivery. The TelemetryService “is used for collecting and reporting metric data within the Redfish service”, with metric report definitions and metric reports; Dell’s GPU Metrics and GPU Statistics are such reports. Older interfaces remain, and Supermicro’s X14/H14 manual lists alerts, SNMP, syslog and SMTP under its notification settings.
For a few servers, one collector per site that subscribes to each BMC’s events and polls the GPU temperature and power resources is enough. Label every series with the server, the slot and the card’s serial number, so that a BMC alert and a DCGM alert for the same card meet in one incident. The BMC should sit on an isolated management network that only the collector and a jump host reach; our guide to securing a GPU server covers that network and BMC firmware.
Setting up both paths on 2 to 4 GPU servers
As an example, take three servers with eight RTX PRO 6000 Server Edition cards each and one with four H200 NVL, 28 GPUs in all. If these were Dell models with the “Datacenter” licence, the GPU Metrics report alone would deliver 40 values per GPU, 1,120 for the site, most of them every 5 seconds, by our arithmetic. This order sets up both paths and tests what each one reports.
- Connect each BMC to the management network, update its firmware to the maker’s current release and record the GPU inventory it shows against nvidia-smi -q.
- Read the Processor resources of type GPU on one server and note which EnvironmentMetrics and ProcessorMetrics properties it fills.
- Subscribe the collector to each BMC’s Redfish events, or configure SNMP or syslog where Redfish events are not available.
- Poll GPU temperature and power out-of-band, and set the thresholds from each card’s slowdown temperature and power limit.
- Install dcgm-exporter in-band, ship the kernel log for Xid lines, and add the ECC, row remapping and NVLink fields your cards support.
- In a maintenance window, unload the NVIDIA driver on one server, note which GPU values the BMC still reports and confirm that a test temperature alert from it arrives; then load the driver and confirm that an in-band test alert arrives too.
- Write down which path owns which alert, so that whoever takes the alert knows where to look first.
Our AI servers are assembled and burn-in tested before shipment, with a test report on request. Send us the cards and server count you plan through the form below, and we reply within one business day with a configuration and quote.
What we supply
We build AI servers to order with out-of-band management over IPMI and the BMC, assembled and burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice. The cards in this article come from our GPU range: the H200 NVL with NVLink bridges, the RTX PRO 6000 Server Edition, the L40S and the L4. The operating system, drivers, CUDA and a container runtime are installed on request, and support for the cards is handled by our engineering partner’s service team. If your monitoring relies on a BMC feature that needs a licence, such as Dell’s iDRAC “Datacenter” licence for telemetry, name it in your request.
FAQ
What is out-of-band GPU monitoring?
Can the BMC monitor GPU temperature?
What GPU telemetry does Redfish provide?
Does iDRAC monitor NVIDIA GPUs?
What is NVIDIA SMBPBI?
Are NVIDIA Xid errors visible out-of-band?
Send us the number of GPU servers and cards you plan, the server maker you prefer and the monitoring system that should receive BMC events and DCGM metrics. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day