A second AI agent doesn’t add safety by default. It adds a handoff, and a handoff is a place where a customer record can change between the moment one agent reads it and the moment another agent acts on it. Take a renewal workflow split into a research agent that pulls usage, billing, and ticket history, and an execution agent that updates the CRM and sends the renewal terms. That split looks like good design. In practice, it’s two independent views of a record that’s still live in the systems underneath both of them.
The gap between “found” and “acted on” is where things go wrong
The research agent reads the account: usage is down 40%, there’s an open billing dispute, support flagged churn risk last week. It hands a summary to the execution agent, which drafts a renewal at last year’s rate and sends it. Thirty seconds passed between the read and the write. In that window, the billing dispute got resolved with a credit, or a rep already called the customer and promised a different number, or the account was reassigned.
Neither agent did anything wrong on its own. The research agent reported what was true when it looked. The execution agent acted on what it was told. But the customer record isn’t a snapshot — it’s a moving target, and no one owns the question of whether the snapshot is still valid by the time it’s used. This is the failure mode that doesn’t show up in a demo, because demos don’t run two agents against a record that a third party — a rep, a billing system, a support ticket — is also touching in real time.
Orchestration means someone owns the state between agents, not just the agents
Splitting research and execution into separate agents is a real pattern, and it’s the right call when the two tasks need different grounding — the research agent working across usage data, billing systems, and ticket history, the execution agent working narrowly against the CRM and messaging systems. The problem isn’t the split. It’s treating the handoff as free.
Orchestration is the principle that has to cover this, and it means something specific: the record’s state at the moment of action is checked, not just referenced. If the execution agent is going to write to a customer record, it needs to re-read the fields it’s about to change immediately before acting, not trust a five-minute-old summary handed to it by another agent. If the underlying values have moved, that’s not a coordination detail — it’s a reason to stop and route to a human, not merge stale and fresh data and act on the average.
This is also where governance and identity do real work instead of decorative work. The execution agent should hold its own scoped identity and its own approval policy, separate from the research agent’s, so a compromised or malfunctioning research step can’t silently walk into a write action. And every agent’s output needs a timestamp and a source claim in the audit trail — not just “renewal sent,” but “renewal sent based on a usage read from 14:02, executed at 14:03” — so if the numbers were stale, that’s findable in minutes, not reconstructed from memory two weeks later. We build this coordination logic explicitly into every multi-agent deployment; it’s part of what a diagnostic call is for — mapping where a handoff in your workflow needs a state check before it’s automated, not after something goes wrong.
Sometimes the honest answer is one agent, not two
There’s a version of this workflow where two agents genuinely earn their coordination cost: the research surface is wide and slow-changing — competitive intel, historical patterns, market data — and the execution surface is narrow and needs to move fast on fresh input. There, the split pays for itself.
But when both agents are touching the same fast-changing record in the same short window, as with the renewal example, the coordination overhead is the risk, not the mitigation of it. A single agent that reads the record, checks the fields it cares about, and acts in one pass has one place to get it wrong instead of two, and one identity and audit trail to reason about instead of a relay between them. Specialization is a design choice with a cost. It’s worth paying when the tasks are genuinely different in grounding or judgment — not just because “one agent per responsibility” sounds like good architecture on a slide.
The question to sit with
Look at the next multi-agent workflow on your roadmap and ask where the handoff actually is — the exact moment one agent’s output becomes another agent’s input. Then ask what happens if the record changes in that gap. If nobody can answer that today, adding a second agent won’t fix it. It will just give the problem a second place to hide.