Fine-Tuning
When to use it
Section titled “When to use it”- You have a base model that is close to what you need, and a dataset of examples for your task.
- The job takes hours, not days.
- You want to pay only while the job runs.
If you need to change most of the model’s weights, or train from scratch, see Distributed Training.
Diagram
Section titled “Diagram” Object storage ── dataset.jsonl ──┐ (Enterprise) ▼ ┌──────────────────────────────────────┐ │ Single GPU instance (on-demand) │ │ startup script: │ │ 1. pull base model + dataset │ │ 2. train LoRA / QLoRA adapter │ │ 3. evaluate on held-out set │ │ 4. upload adapter │ │ attached volume: model cache │ └───────────────┬──────────────────────┘ │ adapter + eval report ▼ Object storage ── adapters/<run-id>/ ──► Model InferenceComponents
Section titled “Components”| Component | Velerion service | Responsibility |
|---|---|---|
| GPU instance | On-demand GPU, one GPU | Sized with Choosing a GPU: 48 GB for LoRA on 8B models, 80–96 GB for QLoRA on 70B models. |
| Job definition | Startup script on the launch form | Runs the four steps above unattended. Machine environment variables are prepended to it as export lines, which makes them the natural place for the run ID, base model name and storage endpoint. |
| Model cache | Attached storage volume | Keeps downloaded base models between runs, so each run does not download tens of gigabytes again. |
| Data and results | Object storage | The training dataset in, and the adapter plus an evaluation report out. |
| Monitoring | Observability | VRAM headroom, utilisation and job logs. |
| Cost control | A project per experiment | Compare what each experiment cost against how much it improved the model. |
Workflow
Section titled “Workflow”- Upload the dataset to an object storage bucket.
- Launch a GPU instance with Enable observability on, the model cache volume attached, and the job as a startup script.
- Follow progress in the host’s logs and VRAM in its host page.
- When the adapter is uploaded, delete the instance. Billing is hourly, and the unused part of the current hour is refunded.
Failure modes
Section titled “Failure modes”| Failure | Effect | Mitigation |
|---|---|---|
| Out of memory | The job crashes at start or on a long example. | Lower the batch size or sequence length, switch LoRA to QLoRA, or move up a memory tier. |
| Instance left running after the job | You pay for an idle GPU. | Self-delete from the script, and alert on idle GPUs. |
| Overfitting on a small dataset | The model looks better in training and worse in use. | Always evaluate on a held-out set before promoting an adapter. |
Trade-offs
Section titled “Trade-offs”- LoRA vs QLoRA: QLoRA fits much larger models on the same GPU but trains more slowly and can lose a little quality. Use LoRA when the model fits in 16-bit.
- Adapters vs merged weights: serving the base model plus a separate adapter lets one inference deployment serve many fine-tunes; merging produces a single, slightly faster model per fine-tune.
