What is AI inference? A practical guide for engineering teams
AI inference is the act of running a trained model on new input to produce an output. For large language models, that means taking a prompt and generating a response token by token. Training builds the model once. Inference runs it every time anyone uses it, which is why it dominates the cost of AI in production.
Inference versus training
| Training | Inference | |
|---|---|---|
| Goal | Learn the weights | Use the weights |
| How often | Once, or occasionally | Every request |
| Bottleneck | Compute | Memory bandwidth and capacity |
| Data risk | Your dataset | Every prompt and document users send |
Fine-tuning, such as LoRA, sits between the two. It is a short training run on your data that adapts an existing model, which is then served by inference like any other.
How LLM inference works, step by step
- Tokenise. The prompt is split into tokens, roughly three quarters of a word each.
- Prefill. The model processes every prompt token in parallel and stores each layer's attention keys and values. This stored state is the KV cache.
- Decode. The model generates one token at a time. Each step reads the whole KV cache and appends to it.
- Stream. Tokens are sent back as they are produced, which is what makes answers appear to type.
The metrics that matter
- Time to first token (TTFT): how long prefill takes. It dominates perceived latency for long prompts.
- Tokens per second: decode speed per request. It decides how fast the answer streams.
- Throughput: total tokens per second across all users. It decides cost per token.
- Context length: how many tokens a request can hold. It is limited by memory, mostly the KV cache.
Why memory, not compute, is the limit
The KV cache grows with every token. For a 70-billion-parameter model with grouped-query attention (80 layers, 8 KV heads of dimension 128, 16-bit values), each token needs about 320 KB. A 128,000-token context is therefore around 40 GB, half an 80 GB GPU, for a single request. That is why long context is expensive, and why techniques like paging, prefix reuse and tiering the cache to host memory and disk matter so much.
Ways to run inference
| Option | You control | Trade-off |
|---|---|---|
| Closed model API | Prompts only | Easiest, but data leaves your control and models change under you |
| Shared open-model API | Model choice | Cheap per token, shared hardware and caches |
| Dedicated GPU endpoint | Model, hardware, region | Isolation and predictability, paid by the hour |
| Self-hosted | Everything | Full control, full operational burden |
kvrun is a dedicated GPU endpoint without the burden: pick an open model and it sizes the GPU, serves it behind an OpenAI-compatible API, and governs the KV cache. See how governed inference works.
Questions
- Is inference cheaper than training?
- Per run, yes. In total, usually not: a popular model is trained once and served millions of times.
- What hardware is used for inference?
- Data-centre GPUs such as NVIDIA A100 and H100 are standard for LLMs, because of their memory capacity and bandwidth.