Agent A Asks Agent B, Agent B Asks Agent A: How to Prevent Recursive Delegation
Agent A asks Agent B to do a task, and Agent B asks Agent A to do another task. How do you prevent recursive delegation? It is a common question in agentic AI design reviews, and it has a clear answer.
Carry a delegation chain with every call: the list of agents, and the job each was asked for, that led to this call. Before a call goes out, refuse it if the target is already in the chain (a cycle) or if the chain is already too long (a depth limit). Let the component that routes the calls own the chain and sign it, so no agent can forge or trim it, and refuse any call whose chain fails that check. Give each branch of parallel work its own chain, because parallel calls are not recursion. And give each chain a budget: a maximum depth, a short lifetime, and a cap on hops.
This stops the loop on its first bounce, before the bounce costs anything. The check needs no shared state, so it gives the same answer on every server that sees the call. The rest of this post explains why the usual fixes fall short, then walks through one example the way the DVARA AI governance platform handles it.
The problem: two good agents make one loop
Take two agents in a travel workspace. Planner builds trip plans. Budget prices them.
A user asks Planner for "a 3-day Lisbon trip under €1,500". Planner asks Budget to price its draft. Budget needs the day-by-day itinerary, so it asks Planner. Planner, still waiting on the price, asks Budget again.
Neither agent is broken. Each does the reasonable thing with what it knows. Together they loop, and every bounce is another round of model calls on your bill.
This is not a rare edge case. It appears as soon as agents can call each other freely, and it gets worse as you add agents: Planner → Budget → FX → Tax → Planner is the same loop with more steps.
Why the usual fixes fall short
Timeouts end the loop only after it has cost money. A 60-second timeout on a loop that bounces every two seconds allows about 30 bounces. Each one is billed.
Loop detectors and rate limits react late. They watch a session and wait for a pattern to repeat enough times to be sure. Here is the Planner and Budget loop under a detector that needs a short cycle to repeat three times:
| Hop | Call | Skill | Result |
|---|---|---|---|
| 1 | user → Planner | plan | allowed |
| 2 | Planner → Budget | price | allowed |
| 3 | Budget → Planner | itinerary | allowed |
| 4 | Planner → Budget | price | allowed |
| 5 | Budget → Planner | itinerary | allowed |
| 6 | Planner → Budget | price | allowed |
| 7 | Budget → Planner | itinerary | refused |
Six hops go through before the seventh is refused. And if an agent drops the session id, the detector cannot join the hops at all.
A per-session list of agents refuses legal work. You could record every agent a session has touched and refuse a repeat. But Planner may call Budget and Weather at the same time, and later call Budget again for a different trip in the same session. One shared list mixes those branches and refuses calls that are fine.
Trusting an agent to say who it is can be faked. If the check relies on a header in which the caller names itself, any caller can name itself anything, or leave out the part of the chain that would get it refused.
The general answer, rule by rule
These rules work for any system where agents call each other through a component that routes the calls: a gateway, a message bus, or an orchestrator.
- Put the chain on the call. Each call carries the path that led to it:
planner → budget. The agent that receives it hands it back on its own onward calls, the way it already forwards a trace id. - Refuse a cycle. If the target is already in the chain, refuse the call. You can choose how strict to be. "Same agent" refuses any return to an agent. "Same agent and same job" allows Planner to be asked for a draft and later for a check, and refuses Planner being asked for the same thing twice.
- Refuse a chain that is too deep. Even with no repeat, a chain of twenty agents is almost always a bug. A depth limit catches loops that change shape, such as an agent that renames its request on each pass.
- Let the router own and sign the chain. The component that forwards the call writes the chain and signs it. An agent can carry the chain but cannot change it. A chain that was changed, has expired, or belongs to another workspace or session is refused. Fail closed: a bad chain is never treated as "no chain".
- Keep parallel branches separate. Because the chain travels with the call, each branch has its own.
planner → budgetandplanner → weathernever see each other, so parallel work is never mistaken for recursion. - Give each chain a budget. A maximum depth, a short lifetime from the first call, and a cap on hops per workspace. A chain that outlives its budget stops, whatever the agents do.
- Answer with a status that says "do not retry". A cycle fails the same way every time. Return a conflict, not a "slow down" error, so the calling agent answers with what it has instead of trying again.
The key design choice is rule 1 together with rule 5: the chain lives on the call, not in a shared store. That is what makes parallel work legal, and it means any server that holds the signing key can check any call without asking another server.
How DVARA does it
DVARA governs each agent-to-agent (A2A) call as a hop through its A2A Gateway. Since 1.8.3, a delegation guard on that gateway applies the rules above.
The gateway signs a delegation chain and forwards it with each hop, in a header called X-DVARA-Delegation. The agent returns it on its next call, together with its session id. Before the gateway forwards a hop, the guard checks the chain for a cycle and for depth. A refused hop never reaches the peer, is never billed, and does not count toward the workspace's hop cap.
Here is the same Planner and Budget example with the guard on:
| Hop | Call | Chain on arrival | Default rule (same agent and skill) | Strict rule (same agent) |
|---|---|---|---|---|
| 1 | user → Planner plan | none | allowed | allowed |
| 2 | Planner → Budget price | planner | allowed | allowed |
| 3 | Budget → Planner itinerary | planner → budget | allowed | refused |
| 4 | Planner → Budget price | planner → budget → planner | refused | never sent |
With the strict rule, Budget receives this on hop 3, and Planner is never called:
409 a2a_delegation_cycle
Delegation cycle: agent planner is already in chain planner → budget.
Budget then prices the draft it already has and answers Planner. The loop ends at its first bounce instead of its seventh hop.
The default rule, same agent and same skill, lets Budget ask Planner for the itinerary once, which is a new job, and refuses the repeat on hop 4. The default depth limit is 8. Each workspace picks its own rule and depth.
Roll it out with observe first. In observe mode the guard records each call it would refuse, and lets the call through. Run it for a week, read the would-refuse events and the depth metric, and then switch to enforce. A chain that fails the signature check is refused even in observe mode, because a bad chain fails closed.
The caller's identity comes from its API key, not from the chain. Bind each agent's key to its agent. The chain says where a call sits in a delegation; the key says who is making it. In 1.8.4 the guard was hardened around this: a chain names the calling agent only through that agent's own key, the gateway takes the chain out of every reply before returning it, and the chain is never written to a log.
What stays as a backstop
The guard is the first line, not the only one. These controls stay on behind it:
- the session loop detector, for agents that drop the chain but keep a session;
- the session kill switch;
- the workspace's monthly A2A hop cap; and
- A2A policies and human approvals.
The guard has limits, and the docs state them. The default rule trusts the skill name the caller sends, so an agent that renames its skill on every pass is stopped by the depth limit, not the cycle rule. If your agents build skill names on the fly, use the strict rule and a lower depth.
Try it
- Stop recursive delegation between AI agents covers the settings, the error codes, and what an agent builder must forward.
- Govern peer-agent calls with the A2A Gateway covers registration, policy, PII and approvals on each hop.