Your KV cache holds your prompts. Here is how to govern it.

· 2 min read · kvrun team

The KV cache is the attention state a language model builds from your prompt and keeps in GPU memory while it answers. It is not an anonymised summary. It is a working representation of everything the user sent, often tens of gigabytes of it, and in most serving stacks it is the least protected copy of your data.

Three ways the KV cache becomes a data risk

1. Sharing between tenants

Prefix caching speeds up serving by reusing cached blocks when two requests start with the same tokens, for example the same system prompt. If those blocks are shared across customers, one tenant's request can reuse state created by another. Published research has shown that timing differences from shared caches can leak information about other users' prompts.

2. No encryption in memory

Traffic is encrypted in transit and disks at rest, but the cache usually sits in plain form in GPU memory, host memory and, once offloaded, on local disk. Anyone with access to that memory or those files sees the context directly.

3. Deletion you cannot prove

When a user exercises the right to erasure, "the cache will be overwritten eventually" is not an answer you can give a regulator. Copies may persist in offload tiers long after the request finished.

What governing the KV cache looks like

RiskControl
Cross-tenant reuseSeparate key hierarchies per tenant; cached blocks are never shared across tenants
Plain memoryEach block sealed with authenticated encryption (AES-256-GCM) under a per-session key
Unprovable deletionErase by destroying the session key; every copy on every tier becomes unreadable at once, and a certificate is issued
No recordAllocation, access, movement and erasure written to an append-only audit log
Context limited by VRAMBlocks tiered from GPU to host memory, disk and object storage, sealed the whole way

Key destruction is the important idea. You cannot reliably find and overwrite every copy of a block. You can destroy the one key that decrypts all of them.

Does this cost performance?

Modern GPUs and CPUs encrypt at many gigabytes per second, so sealing blocks adds little. kvrun's early internal benchmarks show under 3% added time to first token and under 2% lower throughput across common serving engines. Tiering usually more than pays that back by fitting far longer contexts on the same GPU.

Questions to ask your inference provider

  • Is prefix or semantic caching shared between customers?
  • Is cached context encrypted in GPU memory, host memory and on disk?
  • How do you erase one user's context, and what proof do you give?
  • Is access to cached context logged, and can we export the log?

kvrun's answers are on the security model and the governed inference page.