Resources
Everything you need to run open models under governance.
Guides and explainers
From the blog.
- Fine-tuning open models on private data, in the UKHow to fine-tune open LLMs with LoRA and QLoRA on sensitive data without it leaving your control, and how to evaluate and serve the result.
- Private LLM hosting in the UK: dedicated GPUs vs shared APIsComparing ways to host open-source LLMs in the UK: shared APIs, dedicated GPU endpoints and self-hosting. Cost, isolation, residency and effort.
- UK GDPR and LLMs: a checklist for hosting open modelsA practical UK GDPR checklist for running LLM inference and fine-tuning: roles, lawful basis, DPIAs, transfers, security, retention and the right to erasure.
- UK-based AI inference: what it means and why it mattersWhat UK-based AI inference really requires: UK compute, UK storage, a UK operator, and no hidden transfers. A practical guide for UK teams.
- What is AI inference? A practical guide for engineering teamsAI inference is running a trained model to get an answer. How LLM inference works, prefill and decode, the KV cache, latency, throughput and cost.
- What is governed inference?Governed inference means every model request runs under enforceable policy: who can call it, where data lives, what is kept, how it is erased, and the evidence.
- Your KV cache holds your prompts. Here is how to govern it.The KV cache is a working copy of every prompt in GPU memory. The risks of shared prefix caches, and how encryption, isolation and erasure close them.
Trust and product