Govern AI cost with four practical pillars
An AI bill tells you what a provider charged. It does not tell you which product or customer caused the spend, whether the request was worth it, or what should change next. Cost governance connects those questions without turning every engineering decision into a finance project.
Use four pillars: meter and price, attribute and allocate, control and forecast, then optimize and verify. They work in that order. A cheaper route is not useful if the price data is wrong or the answer quality falls.
Managed prices, durable cost records, budgets, forecasts, anomaly detection, chargeback, and cost-aware routing are not included in DVARA Open Source.
Open Source still exposes usage metrics and local audit evidence, and it can reduce provider calls with its exact in-memory cache and standard routing.
Use the four pillars in order
| Pillar | Question it answers | DVARA surface | Check before moving on |
|---|---|---|---|
| Meter and price | What did this call consume, and what should that usage cost? | Token usage, model pricing, cost records | Every active model and provider has a current price; calls with tokens do not appear as unpriced spend |
| Attribute and allocate | Which customer, application, key, model, provider, or cost centre caused it? | Workspaces, opaque API-key IDs, request tags, chargeback | A sample cost row joins back to the workload that sent it |
| Control and forecast | Are we approaching a limit, and what happens when we cross it? | Soft and hard budgets, per-call ceiling, forecasts, anomalies, webhooks | Owners know which cap wins and who receives a soft-limit alert |
| Optimize and verify | Which change reduces spend without making the service worse? | Caching, token controls, routing, canaries, evals | Cost falls while quality, latency, errors, and governance outcomes stay within the agreed range |
Do not start with optimization. First make one ordinary request and reconcile its token counts and cost. Then check that the same request carries the workspace, API-key ID, model, provider, and tags you expect. Only after that should you rely on a budget, forecast, or savings report.
Know how one call becomes a cost
The provider reports usage after a successful model call. DVARA matches the model and provider to the price that is in effect, calculates the token cost, and records the result with its attribution fields.
A price is chosen in this order: an exact model match beats a glob; a more specific glob beats a broader one; and the later effective date breaks a tie. Future prices stay dormant until their effective date. Cached input, cache writes, and reasoning tokens use their own rates when you set them. An empty special rate falls back to the ordinary input or output rate.
Three boundaries matter when you read the totals:
- A call with usage but no matching price remains visible in token usage, but it has no cost record. Zero recorded spend is not proof that the provider call was free.
- A gateway cache hit counts as served usage but creates no upstream cost record because no provider call was made.
- Retries and provider fallback produce at most one usage row and one cost record, attributed to the provider that returned the answer.
See Attribute and reconcile AI spend before handing a number to finance.
Put each control in the right place
Budgets stop or flag spend; they do not reduce the price of a successful call. A soft limit writes an event that you can send to a webhook. A hard limit refuses a request once the applicable spend limit is exhausted. Global, workspace, API-key, and tag-scoped caps can overlap, and the tightest applicable cap wins.
Optimization changes the request path. It can reduce provider calls through caching, limit output tokens, or choose a different model or provider. Those changes need quality and latency checks alongside the cost result. Work through Optimize AI cost safely one lever at a time.
Use the existing task and reference pages
- Cost Management is the canonical guide to Flightdeck pricing, token usage, budgets, forecasts, anomalies, chargeback, and compliance-report screens.
- Attribute cost per workspace is the runnable workspace-to-chargeback workflow.
- Flightdeck budgets and chargeback explains what a workspace team can inspect and tighten for itself.
- Analytics shows where cost data appears alongside metrics, traces, audit evidence, and downloadable reports.