Observability Overview
Velerion Observability is a separate web app for monitoring GPU fleets. It shows the health, utilisation, power and errors of every GPU host that runs the Velerion agent, together with the workloads, metrics and logs coming from those hosts. Open it from Observability in the Console sidebar.
Getting hosts to report
Section titled “Getting hosts to report”A host appears in the app once the Velerion agent is running on it.
| Method | How |
|---|---|
| Launch an instance | Turn on Enable observability on the Console launch form. The agent is installed and set up for you. |
| Connect a cluster | Provision a managed cluster, with Kubernetes, agents and routing handled for you. Coming soon. |
Select Add host / New cluster on the dashboard to open these options at any time. Once a host is registered, its event timeline records when the agent started and when the inventory was registered.
Next steps
Section titled “Next steps”- Fleet & Hosts — the dashboard, host detail and workloads.
- Metrics & Logs — query metrics and search log streams.
- Alerts — get notified before a GPU problem becomes an outage.
- Choosing a GPU — use what you observe to right-size.
