Model Inference
When to use it
Section titled “When to use it”- Applications or customers call a model over HTTP.
- Traffic is steady enough that GPUs stay busy, or valuable enough that latency matters.
- You need to control who can call the model and how often.
Diagram
Section titled “Diagram” Clients │ HTTPS ▼ ┌────────────────────┐ │ CDN & WAAP │ rate limits, bot and attack filtering, TLS └─────────┬──────────┘ ▼ ┌────────────────────┐ │ API gateway │ Enterprise container: auth, quotas, routing, request logs │ (scales 1 → N) │ └───┬───────────┬────┘ │ │ token-authenticated HTTP ▼ ▼ ┌────────┐ ┌────────┐ │ GPU #1 │ │ GPU #2 │ On-demand or reserved instances running a model server └───┬────┘ └───┬────┘ (Docker container), weights cached on a volume └─────┬─────┘ ▼ Observability ── utilisation, VRAM, errors ──► alerts by emailComponents
Section titled “Components”| Component | Velerion service | Responsibility |
|---|---|---|
| Edge | CDN & WAAP | Terminates TLS, filters abusive traffic and absorbs spikes before they reach the gateway. |
| API gateway | Enterprise container | Authenticates callers, enforces quotas, and routes to healthy GPU instances. Keep minimum replicas at one or more, so the gateway never cold-starts. |
| Model servers | On-demand GPUs or reserved capacity | Run an inference server as the instance’s Docker container. Size them with Choosing a GPU. |
| Weights | Attached storage and object storage | Object storage holds the canonical weights; a volume caches them so a replacement instance starts quickly. |
| Monitoring | Observability | Per-GPU utilisation, VRAM and errors; alerts when a server goes silent. |
Scaling
Section titled “Scaling”- Vertically: move to a larger memory tier to fit longer contexts or more concurrent requests.
- Horizontally: add GPU instances and register them with the gateway. Several small GPUs usually serve a small model more cheaply than one large GPU.
- Steady base load is cheapest on reserved capacity; handle peaks with on-demand instances on top.
Failure modes
Section titled “Failure modes”| Failure | Effect | Mitigation |
|---|---|---|
| A GPU instance dies | Its in-flight requests fail. | Run at least two model servers; the gateway retries on the other. Alert on Host silent for (down). |
| KV cache exhausted under load | Latency climbs, then requests are rejected. | Cap concurrency per server in the gateway, and watch gpu_memory_pct. |
| Traffic spike or abuse | GPUs saturate and legitimate callers suffer. | Rate-limit at WAAP and enforce per-caller quotas in the gateway. |
| Offer sold out when scaling | You cannot add capacity quickly. | Keep base load on reserved capacity, and allow the gateway to use more than one GPU model. |
Trade-offs
Section titled “Trade-offs”- Gateway on containers vs on the GPU hosts: a separate gateway costs a little more but keeps authentication and routing up while GPU instances are replaced.
- Quantisation: FP8 or INT4 weights fit more concurrent requests per GPU at a small cost in quality. Measure the quality loss on your own evaluation set.
