Skip to main content
Cost attribution & budgets

Understand AI costs.
Set limits for what comes next.

See which workloads account for your model spend and give each one a budget. DVARA’s AI governance platform combines cost attribution with checks before the next call—not just a dashboard after the bill arrives.

See what happens as a budget fills

Illustration: one enabled $100 monthly budget, an 80% soft threshold, and available spend data. Other policies and caps are outside this example.

Choose the recorded spend
  1. $60 recordedSpend available to the check
  2. DVARA governance$80 soft · $100 hard
  3. ContinueDecision for the next request
Continue

The checked spend is below both thresholds. This budget does not reject the next request.

This is not a spend reservation. Concurrent calls, delayed accounting and unavailable spend data can allow an overshoot.

01 · Understand the spend

Start with a workspace, then inspect the records​

Filter costs by workspace, model, provider and period. Use the underlying records to trace amounts back to token counts and the API key attributed to the call.

Flightdeck · seeded cost demo
Flightdeck Cost Dashboard with populated demo totals and provider and model breakdowns.Flightdeck Cost Dashboard with populated demo totals and provider and model breakdowns.
Seeded historical cost records in an isolated demo. These figures illustrate the reporting UI; they are not a provider invoice or evidence of live model calls.

Cost records, budgets and these Flightdeck views are Enterprise capabilities, not part of the Open Source distribution. They require the relevant configuration and data. See capability availability; Pricing covers production terms.

Know where the dollar figure comes from​

A completed model call needs usable token counts and a matching price. DVARA multiplies input and output tokens by their configured rates and records the total in USD.

Only prices already in effect are eligible. For the selected provider, exact model matches take precedence over patterns; more specific patterns take precedence over broader ones. Keep rates current and check models that do not produce records.

MCP tool calls and A2A hops do not become token-priced model costs in this ledger. This is not the total cost of running an agent workflow.

02 · Choose the response

Warn early. Refuse when the checked spend reaches the cap.​

Create daily, weekly or monthly budgets for an API key, workspace or the whole deployment. Any applicable hard breach can reject a call; a soft breach does not override another cap’s hard limit.

Flightdeck · seeded cost demo
Flightdeck budget list showing enabled monthly caps for seeded demo workspaces.Flightdeck budget list showing enabled monthly caps for seeded demo workspaces.
Configured demo budget records. This screen proves the settings are present, not that a live request was refused.

Soft threshold

Each cap has a warning percentage, 80% by default. Reaching it allows the request and can emit a soft-limit event. Configure a webhook if you need delivery to another system.

Hard limit

On the LLM request path, checked spend at or above a cap rejects the next call with HTTP 402. In-flight calls are not undone, and accounting delays can carry spend past the cap.

When spend is unavailable

General evaluation failures allow requests through. The optional closed cold-start setting covers a specific degraded-counter state; it is not a universal fail-closed guarantee.

Review budget setup and scope →

Keep the per-call check separate​

The optional per-call ceiling is off by default. Its estimate uses the request’s output-token limit and a matching output rate. It does not include the prompt’s input cost, so do not treat it as a full worst-case price.

If the token limit or pricing is missing, this estimate cannot provide that protection. Test the request types and provider routes you actually use.

A cheaper model is a deliberate choice

Configured downgrade rules can change the model as budget pressure rises. They are not automatic savings without configuration: check the replacement model’s quality, capabilities and policies before enabling a rule.

Review per-call and downgrade controls →
03 · Review trends and allocate costs

Use projections as planning inputs, not promises​

Forecasts

The dashboard uses recent spend to project additional spend over the remaining days of the month. That projection is not the month-to-date total plus future spend. New workloads can change it quickly.

Anomalies

Checks compare today’s spend with the preceding 30 complete UTC days. Workspace age and available history matter; a workspace younger than that baseline is skipped. Thresholds are configurable.

Cost allocation

Chargeback reports summarize recorded costs for a chosen scope and period, with PDF and CSV outputs. Validate attribution and rate coverage before sharing them; they are not reconciled provider invoices.

Caller-supplied tags can help group costs, but they do not verify who owns the spend. Use controlled workspace and key assignment for accountability.

Before you set the first cap

Is the recorded amount my provider bill?

No. DVARA calculates USD costs from available token usage and configured rates. Provider discounts, extra charges and billing adjustments are not automatically reconciled. Compare the records with your provider invoice.

What happens if a model has no matching price?

The cost calculation writes no cost record for that call. Missing usable usage has the same effect. An absent record is not a free request, and incomplete accounting weakens spend-based controls.

Does a hard limit guarantee an exact spending ceiling?

No. The check uses spend already visible to it; it does not reserve the cost of every concurrent request. General budget-evaluation failures allow requests through. A separately configured closed cold-start posture applies only to its specific degraded-counter condition, not every failure.

Can I allocate costs by team or feature?

Use workspaces and API keys as the primary attribution dimensions. Applications can also send string tags in request metadata. Those tags are caller-declared labels, not independently verified ownership. Shadow model calls use the shadow key label and a shadow tag when usage and pricing allow a cost record; shadow accounting is best effort, so verify coverage before allocating those costs.

Start with one workspace and a small test budget.​

Check the rates, inspect a cost record, then test the warning and refusal paths before rolling out more caps.