Skip to content

Model Inference

  • Applications or customers call a model over HTTP.
  • Traffic is steady enough that GPUs stay busy, or valuable enough that latency matters.
  • You need to control who can call the model and how often.
Clients
│ HTTPS
▼
┌────────────────────┐
│ CDN & WAAP │ rate limits, bot and attack filtering, TLS
└─────────┬──────────┘
▼
┌────────────────────┐
│ API gateway │ Enterprise container: auth, quotas, routing, request logs
│ (scales 1 → N) │
└───┬───────────┬────┘
│ │ token-authenticated HTTP
▼ ▼
┌────────┐ ┌────────┐
│ GPU #1 │ │ GPU #2 │ On-demand or reserved instances running a model server
└───┬────┘ └───┬────┘ (Docker container), weights cached on a volume
└─────┬─────┘
▼
Observability ── utilisation, VRAM, errors ──► alerts by email
Component Velerion service Responsibility
Edge CDN & WAAP Terminates TLS, filters abusive traffic and absorbs spikes before they reach the gateway.
API gateway Enterprise container Authenticates callers, enforces quotas, and routes to healthy GPU instances. Keep minimum replicas at one or more, so the gateway never cold-starts.
Model servers On-demand GPUs or reserved capacity Run an inference server as the instance’s Docker container. Size them with Choosing a GPU.
Weights Attached storage and object storage Object storage holds the canonical weights; a volume caches them so a replacement instance starts quickly.
Monitoring Observability Per-GPU utilisation, VRAM and errors; alerts when a server goes silent.
  • Vertically: move to a larger memory tier to fit longer contexts or more concurrent requests.
  • Horizontally: add GPU instances and register them with the gateway. Several small GPUs usually serve a small model more cheaply than one large GPU.
  • Steady base load is cheapest on reserved capacity; handle peaks with on-demand instances on top.
Failure Effect Mitigation
A GPU instance dies Its in-flight requests fail. Run at least two model servers; the gateway retries on the other. Alert on Host silent for (down).
KV cache exhausted under load Latency climbs, then requests are rejected. Cap concurrency per server in the gateway, and watch gpu_memory_pct.
Traffic spike or abuse GPUs saturate and legitimate callers suffer. Rate-limit at WAAP and enforce per-caller quotas in the gateway.
Offer sold out when scaling You cannot add capacity quickly. Keep base load on reserved capacity, and allow the gateway to use more than one GPU model.
  • Gateway on containers vs on the GPU hosts: a separate gateway costs a little more but keeps authentication and routing up while GPU instances are replaced.
  • Quantisation: FP8 or INT4 weights fit more concurrent requests per GPU at a small cost in quality. Measure the quality loss on your own evaluation set.