Soft threshold
Each cap has a warning percentage, 80% by default. Reaching it allows the request and can emit a soft-limit event. Configure a webhook if you need delivery to another system.
See which workloads account for your model spend and give each one a budget. DVARA’s AI governance platform combines cost attribution with checks before the next call—not just a dashboard after the bill arrives.
Illustration: one enabled $100 monthly budget, an 80% soft threshold, and available spend data. Other policies and caps are outside this example.
The checked spend is below both thresholds. This budget does not reject the next request.
This is not a spend reservation. Concurrent calls, delayed accounting and unavailable spend data can allow an overshoot.
Filter costs by workspace, model, provider and period. Use the underlying records to trace amounts back to token counts and the API key attributed to the call.

Cost records, budgets and these Flightdeck views are Enterprise capabilities, not part of the Open Source distribution. They require the relevant configuration and data. See capability availability; Pricing covers production terms.
A completed model call needs usable token counts and a matching price. DVARA multiplies input and output tokens by their configured rates and records the total in USD.
Only prices already in effect are eligible. For the selected provider, exact model matches take precedence over patterns; more specific patterns take precedence over broader ones. Keep rates current and check models that do not produce records.
MCP tool calls and A2A hops do not become token-priced model costs in this ledger. This is not the total cost of running an agent workflow.
Create daily, weekly or monthly budgets for an API key, workspace or the whole deployment. Any applicable hard breach can reject a call; a soft breach does not override another cap’s hard limit.

Each cap has a warning percentage, 80% by default. Reaching it allows the request and can emit a soft-limit event. Configure a webhook if you need delivery to another system.
On the LLM request path, checked spend at or above a cap rejects the next call with HTTP 402. In-flight calls are not undone, and accounting delays can carry spend past the cap.
General evaluation failures allow requests through. The optional closed cold-start setting covers a specific degraded-counter state; it is not a universal fail-closed guarantee.
The optional per-call ceiling is off by default. Its estimate uses the request’s output-token limit and a matching output rate. It does not include the prompt’s input cost, so do not treat it as a full worst-case price.
If the token limit or pricing is missing, this estimate cannot provide that protection. Test the request types and provider routes you actually use.
Configured downgrade rules can change the model as budget pressure rises. They are not automatic savings without configuration: check the replacement model’s quality, capabilities and policies before enabling a rule.
Review per-call and downgrade controls →The dashboard uses recent spend to project additional spend over the remaining days of the month. That projection is not the month-to-date total plus future spend. New workloads can change it quickly.
Checks compare today’s spend with the preceding 30 complete UTC days. Workspace age and available history matter; a workspace younger than that baseline is skipped. Thresholds are configurable.
Chargeback reports summarize recorded costs for a chosen scope and period, with PDF and CSV outputs. Validate attribution and rate coverage before sharing them; they are not reconciled provider invoices.
Caller-supplied tags can help group costs, but they do not verify who owns the spend. Use controlled workspace and key assignment for accountability.
No. DVARA calculates USD costs from available token usage and configured rates. Provider discounts, extra charges and billing adjustments are not automatically reconciled. Compare the records with your provider invoice.
The cost calculation writes no cost record for that call. Missing usable usage has the same effect. An absent record is not a free request, and incomplete accounting weakens spend-based controls.
No. The check uses spend already visible to it; it does not reserve the cost of every concurrent request. General budget-evaluation failures allow requests through. A separately configured closed cold-start posture applies only to its specific degraded-counter condition, not every failure.
Use workspaces and API keys as the primary attribution dimensions. Applications can also send string tags in request metadata. Those tags are caller-declared labels, not independently verified ownership. Shadow model calls use the shadow key label and a shadow tag when usage and pricing allow a cost record; shadow accounting is best effort, so verify coverage before allocating those costs.
Check the rates, inspect a cost record, then test the warning and refusal paths before rolling out more caps.