Skip to content

Alerts

Alerts notifies a channel when a GPU or host metric crosses a threshold, or when a host goes down. Firing now at the top of the page lists alerts that are active.

  1. Name the rule, for example GPU temperature critical.
  2. Choose a severity: INFO, WARN or CRIT.
  3. Build a condition from a signal, a comparison operator and a threshold, for example GPU temperature > 87 °C. Select + condition to require more than one condition.
  4. Set Sustained for — 0 s, 30 s, 1 min, 5 min, 15 min or 30 min. The condition must hold for this long before the alert fires.
  5. Choose the channel, then select Save rule. Reset clears the form.

Saved rules are listed under Rules with how many are active.

Group Signals
GPU health GPU temperature, GPU power draw, GPU memory bandwidth
GPU capacity GPU utilization (idle), GPU VRAM used
GPU errors GPU ECC uncorrectable (DBE), GPU ECC correctable (SBE), GPU XID error
Host Host CPU utilization, Host memory used, Host disk used, Host load (1m)
Availability Host silent for (down) — fires when a host stops reporting.

Alerts are delivered by email. The addresses under Channels · Email receive notifications; select Add email to add another.

Rule Severity Condition Sustained for
Host down CRIT Host silent for (down) 5 min
Uncorrectable memory error CRIT GPU ECC uncorrectable (DBE) > 0 0 s
XID error WARN GPU XID error > 0 0 s
GPU running hot WARN GPU temperature > 85 °C 5 min
Idle GPU INFO GPU utilization (idle) < 5 % 30 min

The Idle GPU rule is a cost control, not a health check. On-demand GPUs are billed by the hour whether they are busy or not, so an idle alert catches instances that someone forgot to delete.