Alerts
Alerts notifies a channel when a GPU or host metric crosses a threshold, or when a host goes down. Firing now at the top of the page lists alerts that are active.
Creating a rule
Section titled “Creating a rule”- Name the rule, for example GPU temperature critical.
- Choose a severity: INFO, WARN or CRIT.
- Build a condition from a signal, a comparison operator and a threshold, for example GPU temperature > 87 °C. Select + condition to require more than one condition.
- Set Sustained for — 0 s, 30 s, 1 min, 5 min, 15 min or 30 min. The condition must hold for this long before the alert fires.
- Choose the channel, then select Save rule. Reset clears the form.
Saved rules are listed under Rules with how many are active.
Signals
Section titled “Signals”| Group | Signals |
|---|---|
| GPU health | GPU temperature, GPU power draw, GPU memory bandwidth |
| GPU capacity | GPU utilization (idle), GPU VRAM used |
| GPU errors | GPU ECC uncorrectable (DBE), GPU ECC correctable (SBE), GPU XID error |
| Host | Host CPU utilization, Host memory used, Host disk used, Host load (1m) |
| Availability | Host silent for (down) — fires when a host stops reporting. |
Channels
Section titled “Channels”Alerts are delivered by email. The addresses under Channels · Email receive notifications; select Add email to add another.
A starting rule set
Section titled “A starting rule set”| Rule | Severity | Condition | Sustained for |
|---|---|---|---|
| Host down | CRIT | Host silent for (down) | 5 min |
| Uncorrectable memory error | CRIT | GPU ECC uncorrectable (DBE) > 0 | 0 s |
| XID error | WARN | GPU XID error > 0 | 0 s |
| GPU running hot | WARN | GPU temperature > 85 °C | 5 min |
| Idle GPU | INFO | GPU utilization (idle) < 5 % | 30 min |
The Idle GPU rule is a cost control, not a health check. On-demand GPUs are billed by the hour whether they are busy or not, so an idle alert catches instances that someone forgot to delete.
