Skip to content

Observability Overview

Velerion Observability is a separate web app for monitoring GPU fleets. It shows the health, utilisation, power and errors of every GPU host that runs the Velerion agent, together with the workloads, metrics and logs coming from those hosts. Open it from Observability in the Console sidebar.

A host appears in the app once the Velerion agent is running on it.

Method How
Launch an instance Turn on Enable observability on the Console launch form. The agent is installed and set up for you.
Connect a cluster Provision a managed cluster, with Kubernetes, agents and routing handled for you. Coming soon.

Select Add host / New cluster on the dashboard to open these options at any time. Once a host is registered, its event timeline records when the agent started and when the inventory was registered.

  • Fleet & Hosts — the dashboard, host detail and workloads.
  • Metrics & Logs — query metrics and search log streams.
  • Alerts — get notified before a GPU problem becomes an outage.
  • Choosing a GPU — use what you observe to right-size.