Governed inference for the enterprise

Private, auditable inference for every open model.

Deploy and fine-tune open models on dedicated GPUs in the UK, behind a KV cache that is encrypted, isolated per tenant and provably erasable. Longer context on the GPUs you pay for, and the evidence your auditors ask for.

EncryptedTenant isolatedAuditedkvrun · governed kv cache

client

POST /v1/chat/completions
  host    <your-endpoint>
  auth    Bearer ••••••••  scope=acme
  model   qwen3.5-9b

{"role": "user",
 "content": "Summarise the Q3 claims file."}

<- 200  streaming  46ms to first token

governance log

  1. 09:41:07.114kv.allocws=acme session=s_8f2c blocks=512 cipher=aes-256-gcm
  2. 09:41:07.115policyresidency=eu-west tenant_share=deny decision=allow
  3. 09:41:07.161decodemodel=qwen3.5-9b ttft=46ms stream=open
  4. 09:41:08.902kv.tierblocks=128 gpu -> host sealed=true
  5. 09:41:09.020auditevent=decode tokens=1284 appended
  6. 09:44:51.506erasesession=s_8f2c requested_by=dpo@acme
  7. 09:44:51.509erasekeys=destroyed blocks=unreadable certificate=issued

Illustrative session. Timestamps and identifiers are examples.

Per-block KV encryption, with keys per workspace
AES-256
Time-to-first-token overhead for governance
<3%
More context per 80GB GPU, through tiering
4×
Provable erasure of a session, on request
<60s

The platform

From open weights to audited answers.

The KV cache holds 60 to 80% of inference memory on long-context work, and in most stacks it is shared, unencrypted and forgotten. kvrun runs your models on it as a governed asset, so the same deployment is faster, denser and defensible.

See the models ↗
Any open model, one endpoint

01

Private deployments

Pick an open model and get a private, OpenAI-compatible endpoint on dedicated compute. The GPU is sized from the weights, so you never research an instance type.

  • Curated catalog, or your own weights
  • API only, or API plus hosted chat
  • Dedicated GPUs, London region by default
Your data trains your model only

02

Fine-tuning

Hand it a dataset, get a tuned model back. One LoRA or QLoRA job that starts, finishes and shuts itself down, then serves in one click.

  • LoRA and QLoRA, no training loop
  • Evaluated on every run
  • Your data trains your model only
Every request checked, sealed, logged

03

Governed KV cache

The memory that holds your context, treated as data: encrypted, isolated per tenant, tiered beyond VRAM and logged on every access.

  • Encrypted and isolated per tenant
  • Tiered from GPU to object storage
  • Semantic reuse, not just prefix hits

Performance

Governance that stays out of the hot path.

Encryption, policy checks and audit writes add under 3% to time to first token and under 2% to throughput, across vLLM, SGLang, TensorRT-LLM and llama.cpp. Tiering the cache from GPU memory to host, disk and object storage is what buys the headroom back.

Early internal benchmarks against a stock serving baseline. Your workload will differ.

KV cache tiered from GPU to storage

Context per GPU

kvrun128K+
Baseline32K

Concurrent tenants

kvrun50+
Baseline10

GPU utilisation

kvrun75%+
Baseline45%

Security and privacy

Your context is not our product.

Memory overwrite is not erasure, and a shared prefix cache is not isolation. kvrun treats every block of context as data it is accountable for.

Isolated

Your context. Your keys.

Separate key hierarchies per tenant. Prefix caching never hands one customer's block to another.

Erasable

Forget on command.

Destroy a session's key and every block it sealed is unreadable, everywhere it was stored, with a certificate to prove it.

Audited

Memory with a paper trail.

Allocation, access, movement and erasure land in an append-only log you can export.

Encrypted
Every KV block is sealed with AES-256-GCM under a key unique to your workspace and session. GPU memory holds ciphertext, not your prompts.
Resident
Policy-as-code decides where context may live. Policies ship through Git, not through engine rebuilds.
Training
Nothing you send trains anything for anyone else. Your dataset touches your model and stops there.
KV cache blocks per workspace across four storage tiers, before and after a session is erased
WorkspaceGPUHostDiskObject
acmekey a18 blocks6 blocks4 blocks3 blocks
globexkey g75 blocks7 blocks3 blocks5 blocks
acme, erasedkey destroyed8 blocks6 blocks4 blocks3 blocks
Rows never share a block, and blocks stay sealed as they move from GPU to object storage. Destroy a session's key and every copy, on every tier, goes dark at once.

Open models · one console

Deploy privately.

A curated catalog, every entry sized against the hardware we run so nothing loads and falls over. Need something else? Bring your own weights from Hugging Face.

Model catalog with type, parameter count and context window
ModelTypeParamsContext
Qwenqwen3.6-35b-a3bText36B · 3B active256K
Qwenqwen3.6-27bText27.8B256K
Qwenqwen3.5-9bText9.7B256K
Qwenqwen3.5-4bText4.7B128K
DeepSeekdeepseek-r1-distill-32bText32.8B128K
Googlegemma-3-27bText27.4B128K
Mistralmistral-small-3.1-24bText24B128K
Qwenqwen3-coder-30b-a3bCode30.5B · 3B active256K
Mistraldevstral-small-24bCode24B128K
Qwenqwen2.5-coder-32b-awqCode32.8B128K
Qwenqwen3.8-27bVision27.8B256K
Qwenqwen2.5-vl-32bVision32B128K
Googlegemma-3-12b-visionVision12B128K

The accelerator is chosen from the weights at deploy time, and the exact monthly cost appears before you confirm.

Or start with a few lines of code.

Every deployment is an OpenAI-compatible endpoint. Change the base URL and the key, and keep the code you already have.

  • OpenAI-compatible, so existing SDKs just work
  • Scoped API keys, locked to a workspace
  • Encrypted, isolated KV cache on every request
  • Billed by the hour, never by the seat
curl https://<your-endpoint>/v1/chat/completions \
  -H "Authorization: Bearer $KVRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5-9b",
    "messages": [{ "role": "user", "content": "Hello, kvrun!" }],
    "stream": true
  }'

Compliance

Audit prep in one click, not two weeks.

The evidence is produced as the system runs, so there is nothing to reconstruct later. Erasure certificates, isolation defaults and access logs map onto the controls your auditors test.

Fine-tunes carry their own record. Every run holds out a split, scores the model against it, and writes a verdict mapped to the EU AI Act articles it answers to, kept beside the weights.

GDPR, SOC 2, ISO 27001 evidence
Regulations and the evidence kvrun produces for each
FrameworkRequirementEvidence
GDPR Art. 17Right to erasureErasure certificate per session
GDPR Art. 25Privacy by designEncryption and isolation on by default
SOC 2 CC6.1Access controlPer-tenant key isolation
SOC 2 CC7.2MonitoringContinuous audit log
ISO 27001Media disposalCryptographic erasure
EU AI ActEvaluation and oversightEval record on every fine-tune

Pricing

No seats. No plan to outgrow.

Pay for the hours you run

Accelerator hours and storage, billed only while a deployment is up. Tear it down whenever you like.

Fine-tunes charged once

A training run is billed for the time it takes, then the hardware goes away.

The price you approve

The monthly estimate appears before you confirm, and it is the price on the invoice.

Governed, private, open models

Any open model. Full speed. Your context stays yours.