Stats

List GPU device statistics. Displays a snapshot of current performance metrics including utilization, power, frequency, temperature, memory, and error counters.

Synopsis

xpu-smi stats
xpu-smi stats -d [deviceId]
xpu-smi stats --device [deviceId]
xpu-smi stats --device [pciBdfAddress]
xpu-smi stats --device [deviceId] -j
xpu-smi stats --device [deviceId] -e
xpu-smi stats --device [deviceId] -e -j
xpu-smi stats --device [deviceId] -r
xpu-smi stats --device [deviceId] -r -j
xpu-smi stats --device [deviceId] --samples [count] --interval [milliseconds]
xpu-smi stats --device [deviceId] --list-offline-pages

Options

-h, --help

Print this help message and exit.

-j, --json

Print result in JSON format.

-d <deviceId>, --device <deviceId>, --id <deviceId>

The device ID or PCI BDF address to query. Accepts a comma-separated list to query several devices at once (e.g., -d 0,1,4). If omitted, statistics for all detected devices are displayed.

-e, --eu

Show EU (Execution Unit) array metrics: Active, Stall, and Idle percentages.

-r, --ras

Show RAS (Reliability, Availability, Serviceability) error metrics, including correctable and uncorrectable error counters per category.

--list-offline-pages

List offline memory pages. This option is exclusive — no other stats are shown when it is used.

Requires offline-page reporting support from the GPU and its driver. Where it is unavailable the command reports Offline memory page reporting is not supported on this device or driver and exits non-zero; with -j each device object carries "supported": false together with error and error_code fields.

--samples <count>

Number of samples to collect before computing the displayed statistics. Default: 2.

--interval <milliseconds>

Sampling interval in milliseconds between samples. Default: 100.

Output Metrics

The following metrics are reported per device and per tile where applicable:

Metric

Notes

GPU Utilization (%)

Busiest engine on the tile; device average for multi-tile

Compute / Render / Media / Copy Engine Utilization (%)

Per tile; averaged over the engines of that class

EU Array Active / Stall / Idle (%)

Per tile; shown with -e

GPU Power (W)

Per tile and device

GPU Frequency (MHz)

Per tile and device

Media Frequency (MHz)

Per tile and device

GPU Core Temperature (°C)

Per tile

GPU Memory Temperature (°C)

Per tile

GPU VR Temperature (°C)

Per tile (voltage regulator)

Fan Speed (%)

Per fan

GPU Memory Read / Write (kB/s)

Per tile

GPU Memory Bandwidth (%)

Per tile

GPU Memory Used (MiB)

Per tile

PCIe Read / Write (kB/s)

Device level

Fabric RX / TX (kB/s)

Device level; shown when fabric ports present

RAS Error Counters

Per category (Reset, Programming, Driver, Cache, Mem); shown with -r

[domain] Energy Consumed (J)

Device level; energy consumed over the measurement window, summed from the per-tile counters (a delta, not the cumulative counter that xpu-smi dump reports as energy.consumed). The label is prefixed with the power domain the reading came from, e.g. Card Energy Consumed (J); it is left unprefixed when the contributing tiles do not share one domain. N/A when the device exposes no readable subdevice-, card-, package- or GPU-level power domain

Offline Memory Pages

Count; shown with --list-offline-pages

Examples

Show stats for all GPUs:

xpu-smi stats

Show stats for device 0:

xpu-smi stats --device 0

Show stats including EU metrics in JSON format:

xpu-smi stats --device 0 -e -j

Show RAS error counters for device 0:

xpu-smi stats --device 0 -r

List offline memory pages:

xpu-smi stats --device 0 --list-offline-pages

Collect 10 samples at 200ms intervals:

xpu-smi stats --device 0 --samples 10 --interval 200