Skip to main content
Version: Latest (1.9.x dev)

Attribute and reconcile AI spend

Before you send a chargeback report to finance, prove that one real request can be traced from the caller to its token usage and final dollar amount. This catches stale prices, missing attribution, and unpriced models while the numbers are still small enough to inspect by hand.

Enterprise only

Managed pricing, durable cost records, tag-based spend views, and chargeback reports are not included in DVARA Open Source.

Decide what must receive the cost​

Use the narrowest stable identifier that matches how your organization pays for AI:

Attribution levelUse it forWhat DVARA records
WorkspaceCustomer, product, environment, or business unitWorkspace ID on usage, cost, budgets, audit, and reports
API keyOne application or workload inside a workspaceOpaque API-key ID; the bearer secret is not stored with spend
Request tagCost centre, feature, project, campaign, or sessionString-to-string values from metadata.tags; session ID is also available as the session tag
Model and providerModel economics and provider comparisonThe model requested and the provider that returned the answer

Do not use a display name as the only join key. Names change. Workspace IDs, opaque API-key IDs, and deliberately chosen tags give finance and engineering a stable way to discuss the same traffic.

Price every active model and provider​

Create a price for every (provider, model) pair that can answer production traffic. Include separate cached-input, cache-write, and reasoning rates when the provider bills them differently. Set an effective date instead of editing history in place; existing cost records are not recalculated after a pricing change.

After adding a model or fallback provider, look for token rows with no matching cost. DVARA keeps the usage evidence when pricing is missing, but it does not invent a dollar value. Treat an unpriced served model as a reconciliation failure, not as free traffic.

Check one cost by hand​

Suppose a request reports 8,000 input tokens and 800 output tokens. The active price is $2.50 per million input tokens and $10.00 per million output tokens:

PartCalculationCost
Input8,000 ÷ 1,000,000 × $2.50$0.0200
Output800 ÷ 1,000,000 × $10.00$0.0080
Total$0.0200 + $0.0080$0.0280

Open Flightdeck → Platform Flightdeck → Cost Dashboard, filter to the same workspace, model, provider, and time window, and find that request. Then open the Token Usage tab and confirm the input and output counts. The recorded total should match the calculation above before rounding for display.

If the request used prompt caching or reasoning tokens, split those tokens from the ordinary input or output count and apply the configured special rates. See Price cached and reasoning tokens for the exact calculation.

Account for paths that look unusual​

What happenedUsage evidenceCost outcome
Ordinary successful callProvider-reported usageOne cost record when a matching price exists
Streaming callExact usage when the provider sends it; otherwise estimatedOne cost record when pricing exists
Gateway cache hitServed usage marked as a hitNo upstream cost record
Retry or provider fallbackOne final usage rowOne cost record for the provider that answered
Request blocked before dispatchNo provider usageNo cost record
Provider failure with no usageNo provider usageNo cost record
Successful call with no priceUsage remains visibleNo cost record; investigate it as unpriced spend

These rules explain why a DVARA chargeback total and a provider invoice can differ. A provider invoice can also contain taxes, credits, commitments, or traffic that did not pass through DVARA. Reconciliation should explain those differences rather than force the two totals to match.

Reconcile before closing the month​

Run this check for each provider:

  1. Confirm every active model has the intended current price.
  2. Find token usage with no cost and resolve each unpriced model.
  3. Compare provider and model totals for the same time zone and date boundary.
  4. Separate gateway cache hits from upstream calls.
  5. Sample at least one ordinary call, one streaming call, and one fallback when those paths are in use.
  6. Confirm workspaces, API-key IDs, and required business tags are populated.
  7. Generate chargeback only after the checks pass.

Use Attribute cost per workspace for the complete setup and report workflow. Use Cost Management for pricing precedence, filters, exports, and report contents.