Private LLM hosting in the UK: dedicated GPUs vs shared APIs
There are three realistic ways to run an open-source LLM for a UK business: call a shared API, rent a dedicated GPU endpoint, or self-host on your own servers. The right one depends on how sensitive your data is, how steady your traffic is, and how much infrastructure you want to own.
The three options compared
| Shared open-model API | Dedicated GPU endpoint | Self-hosted | |
|---|---|---|---|
| Billing | Per token | Per hour | Capex or reserved servers |
| Isolation | Shared GPUs and caches | Your GPU | Your hardware |
| UK residency | Varies, often not | Choose the region | Wherever you put it |
| Model choice | Provider's list | Any open model that fits | Anything |
| Fine-tuned models | Limited | Yes | Yes |
| Operational effort | None | Low | High: drivers, serving, scaling, patching |
When a dedicated endpoint wins
- You handle personal, privileged or commercially sensitive data and cannot share hardware or caches.
- You need UK residency written into a contract.
- Your traffic is steady enough that an hourly GPU costs less than per-token pricing.
- You want to serve your own fine-tuned model, not just a public one.
Choosing the GPU
The rule of thumb is memory: at 16-bit precision a model needs about 2 GB of GPU memory per billion parameters for its weights, plus room for the KV cache. A 9B model fits comfortably on a 40 GB A100. A 27B model needs around 56 GB, so it runs on an 80 GB H100 or across two A100s. Quantised models (8-bit or 4-bit) need roughly half or a third as much. kvrun makes this decision from the parameter count, so you never pick an instance type.
Keep your existing code
Most tooling speaks the OpenAI API. A dedicated endpoint that is OpenAI-compatible means changing a base URL and a key, not rewriting your application. Every kvrun deployment works this way, with scoped keys locked to a workspace.
What to check before you sign
- Which region runs the GPU, and which regions hold logs and backups?
- Is the GPU dedicated, or shared with other customers?
- Is cached context encrypted and isolated per tenant?
- Can you tear it down at any time, and stop paying immediately?
- Will your data ever train anything for anyone else?
See the models kvrun serves in the model catalog, or read about UK AI inference.