AI inference glossary
The terms that come up when you run language models in production, defined in a sentence or two.
- AI inference
- Running a trained model on new input to produce an output. For language models, generating a response to a prompt, token by token. Read more
- Context window
- The maximum number of tokens a model can consider in one request, prompt and answer together.
- Cryptographic erasure
- Deleting data by destroying the key that decrypts it, so every copy becomes unreadable at once wherever it is stored.
- Data residency
- A requirement that data is stored and processed in a particular country or region, such as the UK.
- Decode
- The second phase of LLM inference, where tokens are generated one at a time, each step reading the full KV cache.
- Dedicated GPU endpoint
- A model served on a GPU reserved for one customer, rather than shared hardware billed per token. Read more
- Governed inference
- Inference where every request runs under enforceable, provable controls: isolation, encryption, residency, erasure and evidence. Read more
- KV cache
- The attention keys and values a transformer stores for every token it has processed, so each new token does not recompute the whole prompt. It grows with context length and is a working copy of the prompt. Read more
- LoRA
- Low-rank adaptation: fine-tuning a small set of adapter weights while the base model stays frozen. Read more
- Open-weight model
- A model whose trained weights are published, so it can be run on your own or rented hardware. Examples include Qwen, Gemma and Mistral.
- OpenAI-compatible API
- An endpoint that accepts the same requests as OpenAI's chat completions API, so existing SDKs work by changing the base URL and key.
- Policy-as-code
- Governance rules written as versioned code and enforced automatically, instead of documents applied by hand.
- Prefill
- The first phase of LLM inference, where the whole prompt is processed in parallel and the KV cache is built.
- Prefix caching
- Reusing cached KV blocks when requests begin with the same tokens. Fast, but a leakage risk if blocks are shared across tenants.
- QLoRA
- LoRA over a base model quantised to 4 bits, so larger models can be fine-tuned on smaller GPUs.
- Restricted transfer
- Under UK GDPR, sending personal data to a country outside the UK. It needs adequacy regulations or safeguards such as the IDTA. Read more
- Tenant isolation
- Guaranteeing that one customer's data, including cached context, can never be read or reused by another.
- Throughput
- Total tokens generated per second across every request on a GPU. It sets the cost per token.
- Time to first token (TTFT)
- How long a request waits before the first token of the answer arrives. Dominated by prefill on long prompts.
- UK-based AI inference
- Inference whose compute, storage and operator are all in the United Kingdom, so prompts and context never make a restricted transfer. Read more