Skip to content

Reference Architectures Overview

Each reference architecture describes a production-shaped deployment on Velerion: the components, how data flows between them, what fails first, and the trade-offs behind the shape. They are starting points to adapt, not blueprints to copy verbatim.

  1. When to use it — the conditions under which this shape is the right answer.
  2. Diagram — the components and the direction of data flow.
  3. Components — what each piece does, and which Velerion service provides it.
  4. Failure modes — what breaks first, and how far the damage spreads.
  5. Trade-offs — what you give up by choosing this shape.
If you need to… Start with
Pick hardware for any GPU workload Choosing a GPU
Pre-train or fully train a model across many GPUs Distributed Training
Adapt an existing model to your data Fine-Tuning
Serve a model behind an API Model Inference
Run a web app with little operational work Web App on Containers
Run a web app with full control over the platform Web App on Kubernetes
Need Service
GPUs by the hour On-Demand GPUs
Guaranteed GPUs for months Reservations
CPU compute Virtual instances, bare metal, Kubernetes, containers
Relational data Managed PostgreSQL
Datasets, checkpoints, uploads Object storage and volumes
Traffic and protection Networking, load balancers, CDN & WAAP
Monitoring and alerts Observability
Cost control Projects and budgets
  • Everything that can be rebuilt is disposable. Instances are cattle. Anything that must survive lives on a volume, in object storage or in a managed database.
  • Every GPU instance reports to Observability. Turn on Enable observability at launch; a GPU you cannot see is a GPU you cannot debug or account for.
  • Every workload has a project. Budgets make overspend visible before the wallet runs dry.
  • Automation uses scoped API keys, never a person’s session. See API keys.