Skip to main content

Self-Hosted LLM Gateway: Governance That Runs Inside Your Perimeter

· 4 min read

For many teams, the deciding question is not only what an AI governance platform does, but where it runs. If prompts contain patient records, financial data, or source code, adding another vendor to the request path can complicate—or prevent—approval.

Self-hosting keeps the enforcement point inside your environment. Policy evaluation and PII handling happen before a request leaves your perimeter, and a denied request does not leave at all.

What "self-hosted" means here​

A self-hosted deployment is one you run inside your own environment — your VPC, your Kubernetes cluster, your on-prem data center — rather than consuming it as a vendor-operated SaaS. Requests from your applications hit your instance. The Gateway still calls out to model providers (unless you're running fully local models), but it does so from inside your perimeter, under your network controls, with your keys, after your governance has run.

Self-hosted and Open Source describe different choices. DVARA Open Source is Apache-2.0 and runs as a file-configured LLM Gateway without a database. The Enterprise platform is also self-hosted, but adds Flightdeck, central persistence, fleet operations, advanced LLM controls, and the MCP and A2A planes. In either case, the deployment and request path stay in your environment.

What self-hosting actually buys you​

Ordered governance-first, because that's the reason to do it:

  • Governance runs before egress. The most sensitive thing you send a model is often the prompt itself. Self-hosting means the request is policy-checked, PII-scanned, and logged before it leaves your network — and a blocked request never leaves at all.
  • Audit stays in your control. Open Source can write a local tamper-evident audit file. Enterprise stores fleet evidence in your database for search, export, and reporting.
  • Data residency you control. You decide which region the platform runs in and which providers it may reach — how you keep EU traffic on EU infrastructure. (This is the bridge to the broader AI compliance and sovereignty story.)
  • Your keys, your vault. Enterprise provider credentials can resolve from your secret store rather than a vendor-operated service.
  • Air-gap-friendly. Paired with local models, it can run with no outbound internet at all.

BYOK: bring your own keys​

A serious self-hosted platform is BYOK — bring your own keys. Each workspace's provider credentials are theirs; the Gateway resolves the right key per request and never forces everyone through one shared secret. Credentials can be stored encrypted at rest, or — better — as references into a vault (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault), so the secret material never lands in the platform's own database at all. Rotate a key, and in-flight resolution picks up the new one without a redeploy.

For the strictest setups, strict BYOK refuses to let a workspace borrow a shared platform key at all — no own credential, no call — so you can guarantee cross-workspace key isolation as a governed policy.

What self-hosting doesn't do for free​

Self-hosting is a governance-and-data-path guarantee, not a magic wand:

  • You operate it. Open Source can run as one process with no database. Enterprise adds PostgreSQL and the services needed for central operations.
  • You still call external providers unless you run local models; self-hosting the platform doesn't self-host OpenAI.
  • You still own the operating work. Self-hosting trades a vendor-managed request path for your own infrastructure, upgrades, and monitoring.

For most regulated buyers, those are easy trades against the alternative of not being allowed to use the product at all.

An evaluation checklist​

  • Does it deploy into your VPC / cluster cleanly (containers, a Helm chart)?
  • Does governance run before egress — policy and PII enforced before the request leaves your network?
  • Does the audit trail stay in your database, outside any vendor's scope?
  • Is it BYOK, with per-workspace credential isolation (and a strict mode)?
  • Can credentials live in your vault as references, not just encrypted in its DB?
  • Can you pin it to a region and restrict which providers it may reach?

Where DVARA fits​

DVARA is an AI governance platform designed to run in your environment. Start with the file-configured Open Source LLM Gateway, or use Enterprise when you need Flightdeck, central evidence, fleet operations, per-workspace credentials, and vault-backed secrets. For the wider regulatory picture, read the AI Audit & Compliance guide.