Skip to main content
Version: Latest (1.9.x dev)

Govern AI cost with four practical pillars

An AI bill tells you what a provider charged. It does not tell you which product or customer caused the spend, whether the request was worth it, or what should change next. Cost governance connects those questions without turning every engineering decision into a finance project.

Use four pillars: meter and price, attribute and allocate, control and forecast, then optimize and verify. They work in that order. A cheaper route is not useful if the price data is wrong or the answer quality falls.

Enterprise only

Managed prices, durable cost records, budgets, forecasts, anomaly detection, chargeback, and cost-aware routing are not included in DVARA Open Source.

Open Source still exposes usage metrics and local audit evidence, and it can reduce provider calls with its exact in-memory cache and standard routing.

Use the four pillars in order​

PillarQuestion it answersDVARA surfaceCheck before moving on
Meter and priceWhat did this call consume, and what should that usage cost?Token usage, model pricing, cost recordsEvery active model and provider has a current price; calls with tokens do not appear as unpriced spend
Attribute and allocateWhich customer, application, key, model, provider, or cost centre caused it?Workspaces, opaque API-key IDs, request tags, chargebackA sample cost row joins back to the workload that sent it
Control and forecastAre we approaching a limit, and what happens when we cross it?Soft and hard budgets, per-call ceiling, forecasts, anomalies, webhooksOwners know which cap wins and who receives a soft-limit alert
Optimize and verifyWhich change reduces spend without making the service worse?Caching, token controls, routing, canaries, evalsCost falls while quality, latency, errors, and governance outcomes stay within the agreed range

Do not start with optimization. First make one ordinary request and reconcile its token counts and cost. Then check that the same request carries the workspace, API-key ID, model, provider, and tags you expect. Only after that should you rely on a budget, forecast, or savings report.

Know how one call becomes a cost​

The provider reports usage after a successful model call. DVARA matches the model and provider to the price that is in effect, calculates the token cost, and records the result with its attribution fields.

A price is chosen in this order: an exact model match beats a glob; a more specific glob beats a broader one; and the later effective date breaks a tie. Future prices stay dormant until their effective date. Cached input, cache writes, and reasoning tokens use their own rates when you set them. An empty special rate falls back to the ordinary input or output rate.

Three boundaries matter when you read the totals:

  • A call with usage but no matching price remains visible in token usage, but it has no cost record. Zero recorded spend is not proof that the provider call was free.
  • A gateway cache hit counts as served usage but creates no upstream cost record because no provider call was made.
  • Retries and provider fallback produce at most one usage row and one cost record, attributed to the provider that returned the answer.

See Attribute and reconcile AI spend before handing a number to finance.

Put each control in the right place​

Budgets stop or flag spend; they do not reduce the price of a successful call. A soft limit writes an event that you can send to a webhook. A hard limit refuses a request once the applicable spend limit is exhausted. Global, workspace, API-key, and tag-scoped caps can overlap, and the tightest applicable cap wins.

Creating or deleting a cap, or changing its period, hard limit, or enabled state in Flightdeck, requires approval from a different budget manager. Approval is bound to the exact cap values; the original requester returns to execute the change before the 15-minute approval expires. This keeps one person from silently removing or weakening a financial control.

Optimization changes the request path. It can reduce provider calls through caching, limit output tokens, or choose a different model or provider. Those changes need quality and latency checks alongside the cost result. Work through Optimize AI cost safely one lever at a time.

Use the existing task and reference pages​

  • Cost Management is the canonical guide to Flightdeck pricing, token usage, budgets, forecasts, anomalies, chargeback, and compliance-report screens.
  • Attribute cost per workspace is the runnable workspace-to-chargeback workflow.
  • Flightdeck budgets and chargeback explains what a workspace team can inspect and tighten for itself.
  • Analytics shows where cost data appears alongside metrics, traces, audit evidence, and downloadable reports.