Skip to main content
Version: Latest (1.9.x dev)

Optimize AI cost without losing quality

The cheapest model is not automatically the cheapest service. A lower model price can be erased by retries, slower responses, poor answers, or extra support work. Optimize one lever at a time and judge the whole result: cost, quality, latency, reliability, and governance.

Start only after pricing and attribution are trustworthy. If they are not, use Attribute and reconcile AI spend first.

Record a baseline you can compare​

Choose one workload and a representative time window. Record these values before changing anything:

MeasureWhy it matters
Cost per successful request or completed taskNormalizes spend when traffic volume changes
Input, output, cached-input, cache-write, and reasoning tokensShows which part of the request is expensive
Quality or eval pass rateStops savings from hiding worse answers
p50 and p95 latencyCatches a cheaper but slower path
Error, retry, and fallback rateShows whether savings depend on unreliable providers
Cache-hit rateSeparates genuine reuse from ordinary provider calls
Policy and guardrail outcomesConfirms the optimized path still applies the required controls

Keep the workload and evaluation set fixed while comparing a change. If both the prompt and the model change at once, you will not know which one caused the result.

Try the low-risk changes first​

Work down this list and stop when the target is met:

  1. Remove unnecessary input. Shorten repeated instructions and avoid sending context the task does not use. Re-run the same eval set.
  2. Set a sensible output limit. Use the request's max_tokens or a Policy as Code token condition. A limit that is too low truncates useful answers, so watch completion quality.
  3. Cache stable, repeatable work. Exact caching is safest when identical requests should return the same answer. It is off by default; enable it only after choosing an entry limit and expiry. Do not cache responses whose freshness or user-specific context makes reuse unsafe.
  4. Canary a cheaper route. Send a small share of real traffic to the candidate, compare cost, quality, latency, and errors, then promote or roll back. Do not replace the primary route in one step.

The response caching guide states the configuration and invalidation boundaries. The canary cookbook shows the complete rollout and rollback workflow.

Use routing to match cost to the task​

Round-robin, weighted, and canary routing let you test provider and model mixes. They are useful when two options meet the same quality bar at different prices, but they do not inspect whether a particular request is simple or complex.

Enterprise only

Cost-aware and intelligent routing are not included in DVARA Open Source.

Cost-aware routing chooses the cheapest eligible provider and can apply a latency constraint. Intelligent routing classifies a request as simple, moderate, or complex and maps that class to a configured model tier. Neither is a substitute for evaluation: verify the selected model on your own tasks and keep a rollback route.

See Routing and load balancing for the exact strategy rules and configuration.

Treat budgets as guardrails, not savings​

Enterprise only

Managed budgets, forecasts, anomaly alerts, and per-call cost ceilings are not included in DVARA Open Source.

A budget tells DVARA when to alert or refuse traffic. It does not make accepted requests cheaper. Use a soft limit early enough for someone to act, connect its event to a webhook, and reserve a hard limit for the point where refusing work is better than spending more.

The optional per-call ceiling rejects a request whose estimated maximum output cost is too high. It is off by default, depends on a matching price and a positive max_tokens, and estimates output cost only. Use it to catch one runaway request, not to predict the final bill.

See Cost Management for cap precedence, defaults, and failure behavior.

Decide whether the change worked​

Compare the candidate with the baseline over the same workload and time window. Promote it only when:

  • cost per successful request or completed task falls;
  • the eval or quality threshold still passes;
  • latency and error rates remain inside the service target;
  • retry and fallback rates do not erase the saving;
  • policy, PII, guardrail, and audit evidence still appear as expected; and
  • finance can still attribute the result to the correct workspace, key, and business tag.

If the cost falls but one of those checks fails, keep the evidence and roll the change back. That is a useful test result, not a failed optimization program.