Choosing a GPU
Picking a GPU is mostly a memory question. Work out how much VRAM the workload needs, then choose the cheapest offer that fits it with headroom, then check whether the interconnect matters.
Step 1 — Estimate VRAM
Section titled “Step 1 — Estimate VRAM”Model weights dominate memory. Multiply the parameter count by the bytes each parameter takes:
| Precision | Bytes per parameter | 8B model | 70B model |
|---|---|---|---|
| FP32 | 4 | 32 GB | 280 GB |
| FP16 / BF16 | 2 | 16 GB | 140 GB |
| FP8 / INT8 | 1 | 8 GB | 70 GB |
| INT4 | 0.5 | 4 GB | 35 GB |
Then add what the workload needs on top of the weights:
| Workload | Rough total VRAM |
|---|---|
| Inference | Weights, plus the KV cache (grows with context length and concurrent requests), plus about 10–20% overhead. |
| LoRA fine-tuning | 16-bit weights, plus small adapter weights and optimiser state, plus activations. |
| QLoRA fine-tuning | 4-bit weights, plus adapters and activations. The cheapest way to tune a large model. |
| Full fine-tuning / training | About 16 bytes per parameter with mixed precision and AdamW (weights, gradients, FP32 master copy, optimiser state), plus activations. An 8B model needs well over 100 GB, so it must be sharded across GPUs. |
Step 2 — Pick a memory tier
Section titled “Step 2 — Pick a memory tier”The GPU marketplace groups naturally by VRAM per GPU:
| VRAM per GPU | GPUs on Velerion | Good for |
|---|---|---|
| 16 GB | RTX A4000, V100 | Small models, development, INT4 inference of 7–8B models. |
| 24 GB | RTX 4090, L4, A10, RTX A5000 | 7–8B inference in 16-bit, LoRA or QLoRA of 7–8B models. L4 is the low-power inference choice. |
| 32–40 GB | V100 32 GB, A100 40 GB | Mid-size models; the A100 adds BF16 and much higher throughput than the V100. |
| 48 GB | RTX A6000, L40, L40S, RTX 6000 Ada | 13B-class inference, 70B inference at INT4, comfortable LoRA of 8B models. L40S is a strong all-rounder. |
| 80 GB | A100 80 GB, H100 | 70B inference across two GPUs, QLoRA of 70B models, the training workhorse tier. |
| 96 GB | RTX Pro 6000, GH200 | Large single-GPU inference and fine-tuning with long contexts. |
| Reservation | H200, B200, B300 and others | Large-scale training and high-throughput serving. See Reservations. |
Multi-GPU offers (for example 2x L40S or 8x RTX4090) multiply the total VRAM, but a model split
across GPUs only benefits if the software shards it (tensor, pipeline or FSDP parallelism).
Step 3 — Check the interconnect
Section titled “Step 3 — Check the interconnect”Each offer card shows PCIe or SXM.
- SXM GPUs are linked by NVLink, which is far faster than PCIe for traffic between GPUs. Choose SXM for multi-GPU training and for tensor-parallel inference of models that do not fit on one GPU.
- PCIe is fine for single-GPU work and for data-parallel jobs that synchronise rarely, such as running several independent replicas of a small model.
Step 4 — Check precision support
Section titled “Step 4 — Check precision support”Older generations lack newer number formats. The V100 has no BF16 support, which most modern training recipes assume. Ada and Hopper GPUs (L4, L40S, RTX 4090, RTX 6000 Ada, H100) support FP8, which roughly halves weight memory and speeds up inference compared with 16-bit.
Quick recommendations
Section titled “Quick recommendations”| Workload | Start with |
|---|---|
| Experimenting, notebooks, small models | A single 24 GB GPU |
| Serving an 8B model | L4 or L40S |
| Serving a 70B model | 2× 80 GB SXM, or one 48 GB GPU at INT4 |
| LoRA on an 8B model | One 48 GB GPU |
| QLoRA on a 70B model | One 80 GB or 96 GB GPU |
| Full fine-tuning of an 8B model | 4–8× 80 GB SXM |
| Pre-training | Reserved H100/H200/B200-class SXM capacity |
Right-size after the first run
Section titled “Right-size after the first run”Launch the smallest offer you think fits, then look at the host page in Observability:
- VRAM near 100% with out-of-memory errors → move up a tier, or quantise.
- GPU utilisation low and SM occupancy low → the GPU is starved by data loading or CPU work; a bigger GPU will not help. Add vCPUs or fix the input pipeline.
- Throttle reasons active → the GPU is running below its rated speed; report it to Support.
