Metrics & Logs
Metrics
Section titled “Metrics”Metrics queries GPU and host metrics. The schema is normalised, so the same metric names work for every provider in your fleet.
Building a query
Section titled “Building a query”- Choose a Metric.
- Choose an Agg (aggregation), for example
avg. - Optionally select + filter to narrow the series.
- Choose how to Group by:
gpu,host,clusterormodel. - Select Run query.
The query box shows the expression the builder produces, for example
avg by (gpu) (gpu_utilization). Data is bucketed automatically for the chosen time range, for
example into 15-second buckets over the last five minutes.
Available metrics
Section titled “Available metrics”| Metric | Measures |
|---|---|
gpu_utilization |
GPU utilisation. |
gpu_memory_used |
VRAM in use. |
gpu_memory_pct |
VRAM in use as a percentage of capacity. |
gpu_temperature |
GPU temperature. |
gpu_power |
GPU power draw. |
sm_occupancy |
Streaming multiprocessor occupancy. |
host_cpu |
Host CPU utilisation. |
host_memory_used |
Host memory in use. |
host_load |
Host load average. |
host_net_rx, host_net_tx |
Host network traffic received and transmitted. |
Logs collects container, systemd and agent streams from across the fleet.
- Stream shows every line. Failures narrows the view to failures.
- Filter by severity: DEBUG, INFO, WARN, ERROR or CRIT. A counter above the list shows how many lines of each level match.
- Sort by Newest first, or change the order.
- Export downloads the matching lines.
If nothing matches, adjust the source, severity, search or time range. To see the logs of a single machine, use the Logs shortcut on its host page.
