The whole plane, on your metal.
Private inference is the full Sideren serving layer deployed inside your VPC or on-prem environment. Prompts, outputs, logs and weights live on hardware you control — nothing egresses, nothing phones home.
A deployment, not a model server.
Six pieces arrive as one plane. Remove any of them and an internal platform team ends up rebuilding it — so none of them are optional.
The endpoint
One address for every deployed model — OpenAI dialect, Anthropic dialect and plain REST answered natively, streaming included.
The models
Any open-weights model from the catalog — text, embeddings, image, video, voice — plus your own fine-tunes and internal weights.
Key custody
Keys are minted, scoped and revoked inside your boundary. Rotation and storage follow your policy, in your vault.
Caps & metering
Per-team ceilings and live usage meters, so internal platform teams can hand out access without handing out risk.
Audit log
Every request accounted for — attributable, exportable, retained on your schedule. It never leaves your storage.
Agent-grade serving
Streaming that survives hour-long loops and concurrency that holds under fleets — the same engine that runs our cloud.
Scoped, deployed, yours.
Scope
Tell us the models, the volume and the environment. Our team responds within 24 hours with a concrete deployment plan.
Deploy
We land the plane with your team — into your VPC or your data centre, sized to your hardware, keyed to your vault.
Operate
Updates arrive as versioned releases you apply on your schedule. Air-gapped environments stay air-gapped.
Built for the strictest room in the building.
Data-residency mandates, regulated workloads, contracts that forbid third-party processing — private inference exists for the traffic that cannot leave. Every guarantee it rests on is structural, and the full posture is written down.