Fine-tuning open models on private data, in the UK
Fine-tuning adapts an existing open model to your domain, tone or task using your own examples. With LoRA and QLoRA it takes minutes to hours on a single GPU, not weeks on a cluster. The harder part is doing it with sensitive data: keeping the dataset in the UK, tied to one model, and evaluated in a way you can show later.
LoRA or QLoRA?
| LoRA | QLoRA | |
|---|---|---|
| What trains | Small adapter matrices; the base stays frozen | The same adapters, over a 4-bit quantised base |
| Memory | Base model at full precision | Much lower, so bigger models fit smaller GPUs |
| Speed | Faster per step | Slower per step, cheaper per run |
| Use when | The model already fits its GPU | You want a larger model than the GPU would otherwise hold |
Keeping private data private
- Minimise first. Remove fields the task does not need, and pseudonymise identifiers.
- One dataset, one model. Your data should train your adapter and nothing else, never a shared base model.
- Short-lived compute. The training job should start, finish and shut down, so nothing lingers on a machine afterwards.
- UK storage. Keep the dataset and the resulting weights in the same UK region as inference.
Evaluate every run, automatically
A fine-tune you cannot explain later is a liability. Hold out a split of the data, score the tuned model against it, and keep the result with the weights. kvrun does this on every run, turning eval loss and perplexity into a plain verdict and writing a record mapped to the EU AI Act articles it answers to.
Serve it
Once training finishes, the tuned model deploys behind the same kind of private, OpenAI-compatible endpoint as any catalog model, with its own governed KV cache.
Related: private LLM hosting in the UK and the UK GDPR checklist.