Governed inference for the enterprise
Private, auditable inference for every open model.
Deploy and fine-tune open models on dedicated GPUs in the UK, behind a KV cache that is encrypted, isolated per tenant and provably erasable. Longer context on the GPUs you pay for, and the evidence your auditors ask for.
EncryptedTenant isolatedAuditedkvrun · governed kv cache
client
POST /v1/chat/completions
host <your-endpoint>
auth Bearer •••••••• scope=acme
model qwen3.5-9b
{"role": "user",
"content": "Summarise the Q3 claims file."}
<- 200 streaming 46ms to first tokengovernance log
- 09:41:07.114kv.allocws=acme session=s_8f2c blocks=512 cipher=aes-256-gcm
- 09:41:07.115policyresidency=eu-west tenant_share=deny decision=allow
- 09:41:07.161decodemodel=qwen3.5-9b ttft=46ms stream=open
- 09:41:08.902kv.tierblocks=128 gpu -> host sealed=true
- 09:41:09.020auditevent=decode tokens=1284 appended
- 09:44:51.506erasesession=s_8f2c requested_by=dpo@acme
- 09:44:51.509erasekeys=destroyed blocks=unreadable certificate=issued
Illustrative session. Timestamps and identifiers are examples.
- Per-block KV encryption, with keys per workspace
- AES-256
- Time-to-first-token overhead for governance
- <3%
- More context per 80GB GPU, through tiering
- 4×
- Provable erasure of a session, on request
- <60s