Skip to content

Fleet & Hosts

Overview is the fleet dashboard. Every figure respects the time range in the top bar.

Tile Shows
Fleet online Hosts reporting, out of the total, and how many are down.
Workloads Running and queued workloads.
Open alerts Firing alerts, split into critical and warning.
Fleet VRAM GPU memory in use across the fleet, out of the total.
Fleet power Power draw across the fleet, against the combined TDP.
MIG instances Multi-Instance GPU partitions, and how many devices they span.

A second row charts the last hour: fleet utilisation, active workloads (split into training and inference), fleet power and average VRAM utilisation. Fleet status mix shows the share of hosts that are up.

Each hexagon is one device, labelled with its model and current utilisation.

  • Group by none, status or cluster.
  • Colour by status (up, down or unknown) or by util.
  • A letter on each hexagon marks the device kind: C CPU, P pod, M MIG, G generic.
  • Zoom controls sit at the bottom right of the map.
Panel Shows
Firing alerts Alerts that are active now. View all opens Alerts.
Top 5 devices A performance summary of the busiest devices in the window.
Top workloads · by util Workloads ranked by SM utilisation, with their cluster and devices.
Devices · inventory Every device with its model, cluster, status, utilisation, VRAM, temperature, power and GPU workloads. Filter by GPU model. Select a device to open its host page.

The host page describes one machine. The header shows its status, ID, GPU model and cluster, with shortcuts to the host’s Logs and to Delete host.

Section Contents
Current Five-minute averages of GPU utilisation, VRAM utilisation and usage, temperature, power against the power limit, and SM occupancy.
System GPU model and count, P-state, driver, CUDA and VBIOS versions, CPU, memory, kernel, uptime and hourly cost.
Metrics Charts of CPU, GPU and VRAM utilisation, VRAM, temperature, power draw, SM occupancy, and the SM, memory and video clocks.
Workloads on host Workloads that reported from this host in the window.
Throttle reasons Whether SW Power Cap, HW Slowdown, HW Power Brake or Thermal throttling is active.
Event timeline Lifecycle events, such as the host booting and the agent starting.
Errors ECC single-bit (corrected) and double-bit (corrected and uncorrected) error counts, and the latest XID code.
GPU processes Processes using the GPU, with PID, user, VRAM, memory share and SM share.
Logs A link to every log line from this host.

Workloads lists the jobs running on your fleet. Workloads appear as soon as runs start reporting.

The summary tiles show the number of workloads (running, idle and stopped), average utilisation, the devices in use, and an Attention count of critical and warning issues. By kind and Status mix break the list down further.

Filter the table by kind — Training, Inference, Batch or Eval — to see each workload’s cluster, devices, runtime, utilisation, trend and status.