Distributed Training
When to use it
Section titled “When to use it”- The model or its optimiser state does not fit on one GPU.
- A job runs for days, so hardware failure during the run is likely rather than possible.
- You train repeatedly and want the same, repeatable environment every time.
For adapting an existing model with LoRA or QLoRA, use Fine-Tuning instead. It is far cheaper.
Diagram
Section titled “Diagram” ┌──────────────────────────────────────────────┐ Object storage │ datasets/ checkpoints/ final-models/ │ (Enterprise, S3-compatible) └──────▲───────────────┬──────────────▲────────┘ │ stream shards │ resume │ upload every N steps ┌──────┴───────────────▼──────────────┴────────┐ │ Training node — 8× GPU, SXM / NVLink │ On-demand or reserved │ ├─ Docker container: trainer (FSDP / DeepSpeed) │ ├─ Attached volume: local shard cache + latest checkpoint │ └─ Velerion agent ──────────────────────────────┼──► Observability └───────────────────────────────────────────────┘ metrics · logs · alertsComponents
Section titled “Components”| Component | Velerion service | Responsibility |
|---|---|---|
| Training node | On-demand GPU (up to 8 GPUs per instance) or reserved capacity | Runs the trainer. Choose SXM offers for NVLink between GPUs. |
| Trainer | Docker container option on the launch form | A pinned image with the framework, CUDA and your code, so every run starts identically. |
| Configuration | Machine environment variables | Hyperparameters, storage endpoints and run IDs. They reach startup scripts, but not Docker containers, so pass container settings through the image or its entrypoint. |
| Local cache | Attached storage volume | Holds dataset shards and the latest checkpoint close to the GPUs. |
| Durable store | Object storage | Datasets, every checkpoint and the final model. Use a dedicated access key for the trainer. |
| Monitoring | Observability | GPU utilisation, SM occupancy, VRAM, throttling, ECC and XID errors per GPU. |
| Cost control | Project with a budget | Tracks what each training run costs. |
Scaling beyond one node
Section titled “Scaling beyond one node”A single on-demand instance tops out at eight GPUs. For multi-node training, request reserved capacity through Reservations and describe the interconnect you need under Additional requirements. Reserved capacity also removes the risk of an offer selling out between runs.
Failure modes
Section titled “Failure modes”| Failure | Effect | Mitigation |
|---|---|---|
| A GPU fails mid-run (uncorrectable ECC, XID) | The job crashes. | Checkpoint to object storage at a fixed interval, and alert on ECC DBE and XID errors with Sustained for 0 s. Resume on a new instance from the latest checkpoint. |
| Data loading cannot keep up | GPUs sit idle while you pay for them. | Watch SM occupancy next to utilisation. Cache shards on the volume and add dataloader workers. |
| Thermal or power throttling | The run slows down silently. | Check Throttle reasons on the host page, and alert on GPU temperature. |
| Wallet runs out | New hours cannot be charged. | Set a project budget and keep a margin in the wallet for the run’s expected length. |
| Offer sells out | You cannot relaunch the same shape. | Reserve capacity for long or repeated training. |
Trade-offs
Section titled “Trade-offs”- On-demand vs reserved: on-demand needs no commitment but can sell out; reservations are cheaper and guaranteed but lock you in for at least three months.
- Checkpoint interval: frequent checkpoints waste GPU time on I/O; rare ones waste GPU time on recomputation after a failure. Aim to lose at most 30–60 minutes of work.
