Skip to main content

Blog

Insights on AI governance, agentic security, and building production-grade AI infrastructure.

Control LLM Costs Before the Invoice

Control LLM Costs Before the Invoice

· 5 min read

The first LLM invoice that makes someone's stomach drop is a rite of passage. Spend is driven by token counts nobody was watching, spread across teams nobody was attributing, on models nobody chose deliberately. By the time it shows up on a bill, the money is gone and the cause is a forensic exercise.

Most tools answer this with a dashboard — visibility after the fact. That's the difference between observation and governance: a dashboard tells you what you spent; a governed control point can stop you from spending it. Real cost governance moves control upstream, to the moment each call is made — attributing it, and enforcing a budget on it, before the request leaves your perimeter. That enforcement lives in the DVARA LLM Gateway, the data plane of the AI governance platform.

DVARA 1.1.0 — Native MCP, SSO Presets, and Turnkey Kubernetes

DVARA 1.1.0 — Native MCP, SSO Presets, and Turnkey Kubernetes

· 4 min read

DVARA 1.1.0 is generally available. It makes three parts of a production deployment easier to use: native MCP protocol support, verified Keycloak and Auth0 SSO presets, and a Kubernetes deployment tested end to end on GKE and DOKS.

Everything below ships in every Enterprise install — the same governance core, the same MCP Proxy, the same agentic controls. 1.1.0 is a drop-in upgrade from 1.0.x.

How LLM Fallback and Failover Work

How LLM Fallback and Failover Work

· 4 min read

Model providers are remarkably good and occasionally unavailable. They rate-limit you at the worst moment, degrade under load, and have real outages. If your application talks to a provider directly, every one of those events is a user-facing failure — and every team ends up writing its own retry loop, usually badly.

Resilience belongs in the same layer as governance, and for the same reason: it can't be trusted if it's scattered and invisible. When failover runs at a single governed control point — the DVARA LLM Gateway, the data plane of the AI governance platform — every retry and provider switch is policy-checked and written to the audit trail, so "we failed over to a different provider" is a recorded, reviewable event, not a silent mystery. LLM fallback and failover move resilience out of application code and onto the governed path.

Multi-Provider LLM Routing: A Practical Guide

Multi-Provider LLM Routing: A Practical Guide

· 7 min read

Multi-provider LLM routing sends each model request to an eligible provider according to a declared strategy while keeping policy, PII controls, rate limits, audit, and cost attribution on the same request path. It solves a platform problem: applications should not each implement their own provider selection and silently drift into different governance behavior.

DVARA is an AI governance platform; the DVARA LLM Gateway is the component that routes model traffic. The application keeps one OpenAI-compatible endpoint, while the Gateway chooses a provider only after governance checks and capability filtering.

How to Manage AI Guardrails as Code

How to Manage AI Guardrails as Code

· 5 min read

Most teams' "AI guardrails" are a few lines in a system prompt ("never reveal PII, ignore instructions in user data") plus a regex or two bolted onto one service. It's better than nothing — until you try to answer basic questions. Which services actually enforce the PII check? What does the injection filter catch, and what does it miss? When did someone last change the toxicity threshold, and what did that change block? A prompt-engineering hack can't answer any of them, and it silently varies from one code path to the next.

Guardrails as code applies the same discipline that policy as code brought to access rules — declarative, version-controlled, centrally enforced — to the content of AI traffic: the prompts going in and the responses coming out.

What Is Policy as Code? A Practical Definition (and Why It Matters for AI)

What Is Policy as Code? A Practical Definition (and Why It Matters for AI)

· 5 min read

Every organization has policies — which resources people can access, what data can leave the building, who has to approve a risky action. The question is never whether you have policies. It's whether anyone can enforce them, test them, or prove what they did. When a policy lives in a document, the answer is usually no.

Policy as code is the practice of writing those rules in a declarative, version-controlled language that a system enforces automatically — so a policy becomes a reviewable, testable, auditable artifact instead of a paragraph in a wiki nobody reads.

LLM Policy as Code: Rules, Enforcement, and Rollout

LLM Policy as Code: Rules, Enforcement, and Rollout

· 8 min read

LLM policy as code expresses model, token, and tool rules in a declarative, version-controlled format and enforces them at a shared control point before a provider receives the request. It replaces scattered application checks with one reviewable answer to “what is allowed, for whom, and what happened when the rule matched?”

DVARA is an AI governance platform; the DVARA LLM Gateway is the component that enforces LLM policy on the request path. The policy is useful because enforcement and evidence travel together—not merely because the rules are written in YAML.

Streaming, Structured Outputs, Observability, and Local Models

Streaming, Structured Outputs, Observability, and Local Models

· 5 min read
Correction, 2026-08-28

This post was published against 1.0.0 and referred to the gateway image as ghcr.io/dvarahq/dvara/dvara-llm-gateway. The package was renamed to ghcr.io/dvarahq/dvara-gateway in 1.7.0, so the original command no longer pulls anything. The commands below have been updated to the current image and tag; the rest of the post is unchanged.

In the previous post, you got DVARA running and sent requests to multiple providers through a single OpenAI SDK client. This post covers the next layer of governance platform capabilities — LLM streaming, structured JSON outputs, Prometheus observability, and local development with Ollama.

Run Your First Governed LLM Call with the OpenAI SDK

Run Your First Governed LLM Call with the OpenAI SDK

· 7 min read
Correction, 2026-08-28

This post was published against 1.0.0 and referred to the gateway image as ghcr.io/dvarahq/dvara/dvara-llm-gateway. The package was renamed to ghcr.io/dvarahq/dvara-gateway in 1.7.0, so the original command no longer pulls anything. The commands below have been updated to the current image and tag; the rest of the post is unchanged.

DVARA is an AI governance platform. Its LLM Gateway accepts the OpenAI API shape, so an existing SDK can use DVARA by changing the base URL. Your application keeps the same client while model requests gain policy, PII checks, budgets, audit, routing, and operational evidence.