LLM Policy as Code: Rules, Enforcement, and Rollout
LLM policy as code expresses model, token, and tool rules in a declarative, version-controlled format and enforces them at a shared control point before a provider receives the request. It replaces scattered application checks with one reviewable answer to “what is allowed, for whom, and what happened when the rule matched?”
DVARA is an AI governance platform; the DVARA LLM Gateway is the component that enforces LLM policy on the request path. The policy is useful because enforcement and evidence travel together—not merely because the rules are written in YAML.
Updated September 24, 2026 for DVARA 1.8.0.
What is LLM policy as code?
LLM policy as code is a way to define AI-traffic violations as machine-readable rules, keep those rules under change control, and evaluate them consistently before model dispatch. A matching rule can deny the call or warn the caller while allowing it to continue.
The implementation has three parts:
- A policy artifact that names the rule, its priority, matching conditions, action, and an operator-written message.
- A central enforcement point that sees AI-specific request attributes such as the model, requested output tokens, and offered tools.
- An evidence record that identifies the policy and rule involved in each decision.
Without central enforcement, a YAML file is only documentation. Without a reviewable artifact, central enforcement becomes opaque application logic. Without evidence, neither one proves what happened on a particular request.
What should policy govern?
Use policy for deterministic access and usage decisions that can be evaluated from request context.
| Question | Policy pattern |
|---|---|
| Which models may this application call? | Model allowlist or denylist |
| How large may an explicitly requested output be? | max_tokens limit |
| Which functions may this application offer to a model? | Tool allowlist or denylist |
| Should a preview model warn rather than block? | WARN_AGENT action |
| Which region, time window, or budget state is allowed? | Enterprise conditions |
| Which MCP server, tool, or argument is allowed? | Enterprise MCP conditions on the MCP plane |
Keep probabilistic content inspection separate. Prompt-injection detection, content filtering, and PII detection are guardrail or data-protection controls; they should not be disguised as deterministic access-policy rules.
Write one policy that fails clearly
This DVARA Open Source policy denies one unapproved model. Put it in gateway.yaml alongside the workspace, route, and API-key fingerprint from the Open Source quickstart:
policies:
- id: approved-models-only
workspace: default
status: ACTIVE
dsl: |
version: "1"
rules:
- id: deny-legacy-model
priority: 10
conditions:
model:
denylist: [mock/legacy]
action: DENY
deny_message: "This model is not approved. Use mock/approved instead."
Restart DVARA Open Source after changing gateway.yaml; the file is read at startup. Then send a request for the denied model:
curl -s -w '\nHTTP %{http_code}\n' \
http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $DVARA_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "mock/legacy",
"messages": [{"role": "user", "content": "Draft a customer reply."}]
}'
The Gateway stops the request before provider dispatch and returns the safe alternative written into the rule:
{
"error": {
"message": "This model is not approved. Use mock/approved instead.",
"type": "policy_violation",
"code": "policy_denied",
"trace_id": "01K4V5R9X21MDN8W5A6E7ZQ3TC"
}
}
HTTP 403
When the local audit file is configured, the same decision is recorded as POLICY_DENIED with the policy ID, rule ID, workspace, reason, and chain-integrity fields. The trace ID changes on every request; use the value from the actual response when investigating a call.
Read conditions as violations
DVARA policy conditions describe when a rule should fire, not the desired steady state. For example, a model allowlist matches when the requested model is outside the list; the DENY action then blocks the violation.
Four rules prevent most policy mistakes:
- Multiple conditions inside one rule use AND. Put independent denial reasons in separate rules.
- Model names are exact and case-sensitive, not prefixes or globs. List every approved dated model name explicitly.
- A
max_tokensrule evaluates an explicit request value. If the request omitsmax_tokens, that condition does not match or inject a default. - Empty conditions are refused rather than accepted as policies that appear active but check nothing.
Write denial messages for the caller, not the policy author. “Use mock/approved instead” is more useful than “rule 10 failed.”
Choose between deny and warn
DENY stops the matching request before provider dispatch. Use it for a control whose violation must not reach the model provider.
WARN_AGENT allows the call to continue and attaches a governance warning to the result and audit trail. Use it when the calling application is expected to react but blocking would be premature, such as a preview-model migration period.
A warning is not a softer spelling of enforcement. If the caller ignores warnings, the control is observational. Define who consumes the warning and what that system does before relying on it.
Separate Open Source enforcement from Enterprise lifecycle
The YAML language is shared, but the available conditions and operating workflow differ.
| Capability | DVARA Open Source | DVARA Enterprise |
|---|---|---|
| Model allowlist/denylist | Yes | Yes |
Explicit max_tokens ceiling | Yes | Yes |
| LLM-request tool allowlist/denylist | Yes | Yes |
| Active policy enforcement | Yes | Yes |
| Draft, Shadow, Archived, dry-run, version history, rollback | No | Yes |
| Data-residency, time-of-day, budget conditions | No | Yes |
| CEL expressions | No | Yes |
| MCP server, tool, and argument conditions | No | Yes, on the MCP plane |
| Central editing and fleet distribution through Flightdeck | No | Yes |
DVARA Open Source 1.8.0 evaluates ACTIVE policies from gateway.yaml. It does not evaluate Shadow policies or provide CEL, data-residency, time-of-day, budget-utilization, or MCP conditions. Unsupported conditions are refused rather than silently ignored.
DVARA Enterprise adds the managed lifecycle: author a Draft, dry-run the fields the test surface can represent, observe representative traffic in Shadow, promote to Active, and roll back by restoring an earlier version as a new revision. Enterprise production rights and support require a licence.
Updated for 1.8.2. This post originally said the licence activates the MCP and A2A planes. In 1.8.2 every feature, the MCP and A2A planes included, runs without a licence for non-production use, within 3 workspaces and 100,000 governed calls a month. A licence removes both limits and gives production rights and support. An expired licence refuses no traffic.
Roll out policy without creating an outage
A reliable rollout is operational, not merely syntactic:
- Name the violation. Write one rule for one independently actionable reason.
- Use representative requests. Include allowed, denied, missing-field, tool-calling, and streaming cases relevant to the rule.
- Check the error contract. The caller needs a stable status, code, trace ID, and useful remediation message.
- Verify evidence. Confirm that the decision names the expected policy, rule, workspace, and reason.
- Bound the rollout. In Enterprise, use Draft, dry-run, and Shadow before Active. In Open Source, test the file against a non-production Gateway before restarting production.
- Plan reversal. Keep the last known-good file or Enterprise version and rehearse the rollback path.
Dry-run and Shadow are not equivalent. Dry-run evaluates a synthetic input and only the fields the test surface supplies. Shadow observes real request context without changing the live decision. Conditions involving tools, time, region, budgets, or MCP context need representative governed traffic rather than assumptions from a simple model-only dry-run.
Policy as code is not application authorization
IAM still determines who can authenticate to applications and administrative surfaces. Application authorization still decides which business operation a user may perform. LLM policy governs the AI request at the shared enforcement point.
These layers should reinforce one another, not substitute for one another. A tool policy can prevent an application from offering cancel_order to a model; it does not decide whether the signed-in human is authorized to cancel order ORD-1048 inside the order service.
Common LLM policy-as-code mistakes
Writing broad rules with unclear messages. Separate model, token, and tool violations so each response tells the caller what to fix.
Assuming an allowlist matches allowed values. It matches the violation—values outside the list—so the associated action can deny them.
Describing Enterprise lifecycle as an Open Source feature. Open Source reads active file-based policy at startup. Draft, Shadow, dry-run, version history, rollback, CEL, and advanced conditions belong to Enterprise.
Forgetting the bypass path. A centrally enforced rule governs only requests that traverse the Gateway. Restrict direct provider credentials and network paths accordingly.
Treating an audit record as proof that storage cannot change. The Open Source audit file is tamper-evident through HMAC signing and hash chaining. Someone with direct filesystem access can still delete data; the mechanism makes changes detectable rather than impossible.
Start with one model rule and one observable result
The fastest useful policy is small: deny one model, send one allowed request and one denied request, then verify the error and audit record. Expand only after the team can explain the behavior and reversal path.
Use the Open Source quickstart for a local governed request, the Policy-as-Code guide for the complete rollout, and the policy condition reference for exact matching behavior. For product-level evaluation, see Policy-as-Code for LLM traffic.