Skip to content

Choosing a GPU

Picking a GPU is mostly a memory question. Work out how much VRAM the workload needs, then choose the cheapest offer that fits it with headroom, then check whether the interconnect matters.

Model weights dominate memory. Multiply the parameter count by the bytes each parameter takes:

Precision Bytes per parameter 8B model 70B model
FP32 4 32 GB 280 GB
FP16 / BF16 2 16 GB 140 GB
FP8 / INT8 1 8 GB 70 GB
INT4 0.5 4 GB 35 GB

Then add what the workload needs on top of the weights:

Workload Rough total VRAM
Inference Weights, plus the KV cache (grows with context length and concurrent requests), plus about 10–20% overhead.
LoRA fine-tuning 16-bit weights, plus small adapter weights and optimiser state, plus activations.
QLoRA fine-tuning 4-bit weights, plus adapters and activations. The cheapest way to tune a large model.
Full fine-tuning / training About 16 bytes per parameter with mixed precision and AdamW (weights, gradients, FP32 master copy, optimiser state), plus activations. An 8B model needs well over 100 GB, so it must be sharded across GPUs.

The GPU marketplace groups naturally by VRAM per GPU:

VRAM per GPU GPUs on Velerion Good for
16 GB RTX A4000, V100 Small models, development, INT4 inference of 7–8B models.
24 GB RTX 4090, L4, A10, RTX A5000 7–8B inference in 16-bit, LoRA or QLoRA of 7–8B models. L4 is the low-power inference choice.
32–40 GB V100 32 GB, A100 40 GB Mid-size models; the A100 adds BF16 and much higher throughput than the V100.
48 GB RTX A6000, L40, L40S, RTX 6000 Ada 13B-class inference, 70B inference at INT4, comfortable LoRA of 8B models. L40S is a strong all-rounder.
80 GB A100 80 GB, H100 70B inference across two GPUs, QLoRA of 70B models, the training workhorse tier.
96 GB RTX Pro 6000, GH200 Large single-GPU inference and fine-tuning with long contexts.
Reservation H200, B200, B300 and others Large-scale training and high-throughput serving. See Reservations.

Multi-GPU offers (for example 2x L40S or 8x RTX4090) multiply the total VRAM, but a model split across GPUs only benefits if the software shards it (tensor, pipeline or FSDP parallelism).

Each offer card shows PCIe or SXM.

  • SXM GPUs are linked by NVLink, which is far faster than PCIe for traffic between GPUs. Choose SXM for multi-GPU training and for tensor-parallel inference of models that do not fit on one GPU.
  • PCIe is fine for single-GPU work and for data-parallel jobs that synchronise rarely, such as running several independent replicas of a small model.

Older generations lack newer number formats. The V100 has no BF16 support, which most modern training recipes assume. Ada and Hopper GPUs (L4, L40S, RTX 4090, RTX 6000 Ada, H100) support FP8, which roughly halves weight memory and speeds up inference compared with 16-bit.

Workload Start with
Experimenting, notebooks, small models A single 24 GB GPU
Serving an 8B model L4 or L40S
Serving a 70B model 2× 80 GB SXM, or one 48 GB GPU at INT4
LoRA on an 8B model One 48 GB GPU
QLoRA on a 70B model One 80 GB or 96 GB GPU
Full fine-tuning of an 8B model 4–8× 80 GB SXM
Pre-training Reserved H100/H200/B200-class SXM capacity

Launch the smallest offer you think fits, then look at the host page in Observability:

  • VRAM near 100% with out-of-memory errors → move up a tier, or quantise.
  • GPU utilisation low and SM occupancy low → the GPU is starved by data loading or CPU work; a bigger GPU will not help. Add vCPUs or fix the input pipeline.
  • Throttle reasons active → the GPU is running below its rated speed; report it to Support.