Attribute and reconcile AI spend
Before you send a chargeback report to finance, prove that one real request can be traced from the caller to its token usage and final dollar amount. This catches stale prices, missing attribution, and unpriced models while the numbers are still small enough to inspect by hand.
Managed pricing, durable cost records, tag-based spend views, and chargeback reports are not included in DVARA Open Source.
Decide what must receive the cost
Use the narrowest stable identifier that matches how your organization pays for AI:
| Attribution level | Use it for | What DVARA records |
|---|---|---|
| Workspace | Customer, product, environment, or business unit | Workspace ID on usage, cost, budgets, audit, and reports |
| API key | One application or workload inside a workspace | Opaque API-key ID; the bearer secret is not stored with spend |
| Request tag | Cost centre, feature, project, campaign, or session | String-to-string values from metadata.tags; session ID is also available as the session tag |
| Model and provider | Model economics and provider comparison | The model requested and the provider that returned the answer |
Do not use a display name as the only join key. Names change. Workspace IDs, opaque API-key IDs, and deliberately chosen tags give finance and engineering a stable way to discuss the same traffic.
Price every active model and provider
Create a price for every (provider, model) pair that can answer production
traffic. Include separate cached-input, cache-write, and reasoning rates when
the provider bills them differently. Set an effective date instead of editing
history in place; existing cost records are not recalculated after a pricing
change.
After adding a model or fallback provider, look for token rows with no matching cost. DVARA keeps the usage evidence when pricing is missing, but it does not invent a dollar value. Treat an unpriced served model as a reconciliation failure, not as free traffic.
Check one cost by hand
Suppose a request reports 8,000 input tokens and 800 output tokens. The active price is $2.50 per million input tokens and $10.00 per million output tokens:
| Part | Calculation | Cost |
|---|---|---|
| Input | 8,000 ÷ 1,000,000 × $2.50 | $0.0200 |
| Output | 800 ÷ 1,000,000 × $10.00 | $0.0080 |
| Total | $0.0200 + $0.0080 | $0.0280 |
Open Flightdeck → Platform Flightdeck → Cost Dashboard, filter to the same workspace, model, provider, and time window, and find that request. Then open the Token Usage tab and confirm the input and output counts. The recorded total should match the calculation above before rounding for display.
If the request used prompt caching or reasoning tokens, split those tokens from the ordinary input or output count and apply the configured special rates. See Price cached and reasoning tokens for the exact calculation.
Account for paths that look unusual
| What happened | Usage evidence | Cost outcome |
|---|---|---|
| Ordinary successful call | Provider-reported usage | One cost record when a matching price exists |
| Streaming call | Exact usage when the provider sends it; otherwise estimated | One cost record when pricing exists |
| Gateway cache hit | Served usage marked as a hit | No upstream cost record |
| Retry or provider fallback | One final usage row | One cost record for the provider that answered |
| Request blocked before dispatch | No provider usage | No cost record |
| Provider failure with no usage | No provider usage | No cost record |
| Successful call with no price | Usage remains visible | No cost record; investigate it as unpriced spend |
These rules explain why a DVARA chargeback total and a provider invoice can differ. A provider invoice can also contain taxes, credits, commitments, or traffic that did not pass through DVARA. Reconciliation should explain those differences rather than force the two totals to match.
Reconcile before closing the month
Run this check for each provider:
- Confirm every active model has the intended current price.
- Find token usage with no cost and resolve each unpriced model.
- Compare provider and model totals for the same time zone and date boundary.
- Separate gateway cache hits from upstream calls.
- Sample at least one ordinary call, one streaming call, and one fallback when those paths are in use.
- Confirm workspaces, API-key IDs, and required business tags are populated.
- Generate chargeback only after the checks pass.
Use Attribute cost per workspace for the complete setup and report workflow. Use Cost Management for pricing precedence, filters, exports, and report contents.