What is AI inference? A practical guide for engineering teams

· 2 min read · kvrun team

AI inference is the act of running a trained model on new input to produce an output. For large language models, that means taking a prompt and generating a response token by token. Training builds the model once. Inference runs it every time anyone uses it, which is why it dominates the cost of AI in production.

Inference versus training

TrainingInference
GoalLearn the weightsUse the weights
How oftenOnce, or occasionallyEvery request
BottleneckComputeMemory bandwidth and capacity
Data riskYour datasetEvery prompt and document users send

Fine-tuning, such as LoRA, sits between the two. It is a short training run on your data that adapts an existing model, which is then served by inference like any other.

How LLM inference works, step by step

  1. Tokenise. The prompt is split into tokens, roughly three quarters of a word each.
  2. Prefill. The model processes every prompt token in parallel and stores each layer's attention keys and values. This stored state is the KV cache.
  3. Decode. The model generates one token at a time. Each step reads the whole KV cache and appends to it.
  4. Stream. Tokens are sent back as they are produced, which is what makes answers appear to type.

The metrics that matter

  • Time to first token (TTFT): how long prefill takes. It dominates perceived latency for long prompts.
  • Tokens per second: decode speed per request. It decides how fast the answer streams.
  • Throughput: total tokens per second across all users. It decides cost per token.
  • Context length: how many tokens a request can hold. It is limited by memory, mostly the KV cache.

Why memory, not compute, is the limit

The KV cache grows with every token. For a 70-billion-parameter model with grouped-query attention (80 layers, 8 KV heads of dimension 128, 16-bit values), each token needs about 320 KB. A 128,000-token context is therefore around 40 GB, half an 80 GB GPU, for a single request. That is why long context is expensive, and why techniques like paging, prefix reuse and tiering the cache to host memory and disk matter so much.

Ways to run inference

OptionYou controlTrade-off
Closed model APIPrompts onlyEasiest, but data leaves your control and models change under you
Shared open-model APIModel choiceCheap per token, shared hardware and caches
Dedicated GPU endpointModel, hardware, regionIsolation and predictability, paid by the hour
Self-hostedEverythingFull control, full operational burden

kvrun is a dedicated GPU endpoint without the burden: pick an open model and it sizes the GPU, serves it behind an OpenAI-compatible API, and governs the KV cache. See how governed inference works.

Questions

Is inference cheaper than training?
Per run, yes. In total, usually not: a popular model is trained once and served millions of times.
What hardware is used for inference?
Data-centre GPUs such as NVIDIA A100 and H100 are standard for LLMs, because of their memory capacity and bandwidth.