Fleet & Hosts
Fleet dashboard
Section titled “Fleet dashboard”Overview is the fleet dashboard. Every figure respects the time range in the top bar.
Summary tiles
Section titled “Summary tiles”| Tile | Shows |
|---|---|
| Fleet online | Hosts reporting, out of the total, and how many are down. |
| Workloads | Running and queued workloads. |
| Open alerts | Firing alerts, split into critical and warning. |
| Fleet VRAM | GPU memory in use across the fleet, out of the total. |
| Fleet power | Power draw across the fleet, against the combined TDP. |
| MIG instances | Multi-Instance GPU partitions, and how many devices they span. |
A second row charts the last hour: fleet utilisation, active workloads (split into training and inference), fleet power and average VRAM utilisation. Fleet status mix shows the share of hosts that are up.
Fleet health map
Section titled “Fleet health map”Each hexagon is one device, labelled with its model and current utilisation.
- Group by none, status or cluster.
- Colour by status (up, down or unknown) or by util.
- A letter on each hexagon marks the device kind: C CPU, P pod, M MIG, G generic.
- Zoom controls sit at the bottom right of the map.
Other panels
Section titled “Other panels”| Panel | Shows |
|---|---|
| Firing alerts | Alerts that are active now. View all opens Alerts. |
| Top 5 devices | A performance summary of the busiest devices in the window. |
| Top workloads · by util | Workloads ranked by SM utilisation, with their cluster and devices. |
| Devices · inventory | Every device with its model, cluster, status, utilisation, VRAM, temperature, power and GPU workloads. Filter by GPU model. Select a device to open its host page. |
Host detail
Section titled “Host detail”The host page describes one machine. The header shows its status, ID, GPU model and cluster, with shortcuts to the host’s Logs and to Delete host.
| Section | Contents |
|---|---|
| Current | Five-minute averages of GPU utilisation, VRAM utilisation and usage, temperature, power against the power limit, and SM occupancy. |
| System | GPU model and count, P-state, driver, CUDA and VBIOS versions, CPU, memory, kernel, uptime and hourly cost. |
| Metrics | Charts of CPU, GPU and VRAM utilisation, VRAM, temperature, power draw, SM occupancy, and the SM, memory and video clocks. |
| Workloads on host | Workloads that reported from this host in the window. |
| Throttle reasons | Whether SW Power Cap, HW Slowdown, HW Power Brake or Thermal throttling is active. |
| Event timeline | Lifecycle events, such as the host booting and the agent starting. |
| Errors | ECC single-bit (corrected) and double-bit (corrected and uncorrected) error counts, and the latest XID code. |
| GPU processes | Processes using the GPU, with PID, user, VRAM, memory share and SM share. |
| Logs | A link to every log line from this host. |
Workloads
Section titled “Workloads”Workloads lists the jobs running on your fleet. Workloads appear as soon as runs start reporting.
The summary tiles show the number of workloads (running, idle and stopped), average utilisation, the devices in use, and an Attention count of critical and warning issues. By kind and Status mix break the list down further.
Filter the table by kind — Training, Inference, Batch or Eval — to see each workload’s cluster, devices, runtime, utilisation, trend and status.
