Multi-Provider LLM Routing in an LLM Gateway
Multi-provider LLM routing sends each model request to an eligible provider according to a declared strategy while keeping policy, PII controls, rate limits, audit, and cost attribution on the same request path. It solves a platform problem: applications should not each implement their own provider selection and silently drift into different governance behavior.
DVARA is an AI governance platform; the DVARA LLM Gateway is the component that routes model traffic. The application keeps one OpenAI-compatible endpoint, while the Gateway chooses a provider only after governance checks and capability filtering.
Updated September 24, 2026 for DVARA 1.8.0.
What is multi-provider LLM routing?
Multi-provider LLM routing is the controlled selection of an upstream model provider for each request. A route matches the requested model, removes providers that cannot satisfy the request, then applies a strategy such as round-robin or weighted distribution to the remaining pool.
That definition has three important boundaries:
- Routing is not model rewriting. The selected provider receives the model identifier from the request. Every provider in a shared route must accept that identifier.
- Routing is not failover. Routing selects the primary provider before dispatch; failover is a recovery step after an eligible provider failure.
- Routing is not governance by itself. Provider selection still needs policy enforcement, PII handling, rate limits, and an evidence trail around it.
Centralizing the decision gives platform teams one place to review routing behavior without moving the decision into every application repository.
How does an LLM routing decision work?
A governed request follows five distinct stages:
- Evaluate governance. Resolve the API key and workspace, then apply policy, guardrails, PII handling, and configured limits.
- Match a route. Compare the request's
modelwith configured route patterns. If no route matches, use the built-in model-prefix mapping. - Filter by capability. Remove providers that cannot satisfy requirements such as structured output, JSON mode, or tool calling.
- Apply the routing strategy. Select one provider from the eligible pool.
- Recover separately. If the upstream call fails and resilience permits recovery, try a compatible registered fallback.
This order prevents a routing decision from bypassing governance. It also prevents failover from sending a structured-output or tool-calling request to a provider that cannot honor it. If no compatible provider exists, the Gateway returns an explicit capability-mismatch error instead of silently degrading the request.
Which routing strategy should you use?
Choose the simplest strategy that expresses the operational decision you need to make.
| Strategy | Use it when | Availability |
|---|---|---|
| Model prefix | A model family has one natural provider, such as gpt-* to OpenAI or claude-* to Anthropic | Open Source and Enterprise |
| Round-robin | Registered providers are interchangeable and should receive an even share | Open Source and Enterprise |
| Weighted | You need a controlled traffic split, such as 80/20, between compatible providers | Open Source and Enterprise |
| Canary | You want to bound exposure to a candidate provider while keeping a baseline | Open Source and Enterprise |
| Latency-aware | Selection should use observed provider latency | Enterprise |
| Cost-aware | Selection should consider configured model cost and an optional latency constraint | Enterprise |
| Geo-aware | Provider region must participate in selection | Enterprise |
| Intelligent | Request complexity should select a configured model tier | Enterprise |
DVARA Open Source 1.8.0 does not provide latency-aware, cost-aware, geo-aware, or intelligent routing. Configuring one of those strategies in the Open Source build is refused at startup rather than silently falling back to another strategy.
Configure weighted routing without changing application code
This gateway.yaml registers OpenAI and Azure OpenAI, then sends roughly 80% of matching traffic to OpenAI and 20% to Azure OpenAI:
providers:
- type: openai
api_key: ${OPENAI_API_KEY}
- type: azure-openai
api_key: ${AZURE_OPENAI_API_KEY}
base_url: ${AZURE_OPENAI_BASE_URL}
routes:
- id: weighted-gpt
model: "gpt*"
strategy: weighted
providers:
- provider: openai
weight: 80
- provider: azure-openai
weight: 20
Both providers receive the same requested model name. The Azure deployment therefore needs to use the model name the application sends; routing does not translate gpt-4o-mini into a different deployment identifier.
After restarting DVARA Open Source with the updated file, the application continues to call the same endpoint:
curl -s http://localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $DVARA_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Summarize this incident in two bullets."}]
}'
The response stays OpenAI-compatible. To verify the routing effect, use the response trace ID to inspect the structured access log; its provider field records which upstream handled the request. Over a meaningful request sample, the observed distribution should approach the configured weights. A handful of calls is not enough to validate a probabilistic split.
Open Source reads routes from gateway.yaml at startup, so changing a route requires a restart. DVARA Enterprise stores routes centrally and distributes updates through Flightdeck without requiring an application redeploy.
Keep capability filtering separate from provider preference
A provider can be healthy and still be wrong for a particular request. Structured outputs, JSON mode, tool calls, streamed tool calls, vision, embeddings, and batch operations differ across providers.
Capability filtering narrows the pool before the strategy runs. This means an 80/20 weighted route is really an 80/20 split across the providers still eligible for that request. If one configured member is unavailable or lacks a required capability, it does not receive traffic merely to preserve the nominal percentage.
This distinction matters when interpreting routing metrics. A changed distribution can reflect health or capability filtering, not a broken weight calculation.
Do not confuse routing with failover
Routing and failover answer different questions:
| Decision | Question | When it runs |
|---|---|---|
| Routing | Which eligible provider should receive this request first? | Before the upstream call |
| Retry | Should the same provider receive another attempt? | After a retryable failure |
| Failover | Can another compatible provider safely take over? | After eligible retries or provider unavailability |
Retries, circuit breakers, timeouts, and fallback are enabled by default in DVARA Open Source, with configuration available per provider. They should still be tested with realistic failures. A fallback is only safe when the alternative provider supports the request and accepts the same model identifier.
For the full recovery path, read LLM fallback and failover.
Put governance before optimization
Provider optimization is useful only if it preserves the controls applied to the request. A centrally routed call should still answer:
- Which workspace and API key initiated it?
- Which policy decision applied before dispatch?
- Was PII logged, blocked, or redacted under the configured action?
- Which provider handled the call, and did a retry or fallback occur?
- What usage and cost were attributed to the request?
DVARA Open Source puts routing, file-configured policy, PII controls, guardrails, rate limits, and an optional local tamper-evident audit file on the same LLM request path. DVARA Enterprise adds Flightdeck, central persistence and fleet operations, advanced Enterprise controls, and the MCP and A2A Gateways.
Common multi-provider routing mistakes
Treating provider names as model mappings. A route selects a provider; it does not translate model identifiers between vendors.
Calling weights cost optimization. A static 80/20 split does not inspect price. Cost-aware selection is a separate Enterprise strategy.
Assuming configuration is live in Open Source. gateway.yaml is read at startup. Restart the Gateway after changing routes or policies.
Ignoring capabilities during failover tests. Test structured output, tool calling, streaming, and vision separately. A successful text fallback does not prove every request shape can fail over.
Letting applications bypass the governed endpoint. Direct provider calls create a second path without the same routing policy or evidence. Restrict provider credentials and network paths so production calls traverse the intended control point.
Start with one route you can verify
Begin with one model family and one explicit operational goal: an even split, a controlled canary, or a static weight. Record the baseline provider distribution, test an unavailable provider, and verify that policy and audit behavior remain intact.
Use the Open Source quickstart to run DVARA locally, then follow the routing and load-balancing reference for every strategy, field, and distribution boundary. For how routing fits the wider governance path, see the DVARA LLM Gateway.