When an AI agent starts doing too many things, the temptation is to split it up. One takes the call and works out what the person wants, another handles billing, another faults and another appointments. It looks like the way a human team is organized, and it looks great in a diagram. Sometimes it’s the right call. Other times it multiplies the points of failure without improving anything.
What a multi-agent system is
A multi-agent system splits a task across several agents, each with its own instructions, its own tools and its own piece of the problem. If you’re not sure what an agent is, we explain it in what it takes to call a chatbot an agent.
Two ways of organizing them come up again and again:
- Orchestrator and workers. A lead agent receives the task, breaks it down, hands the pieces to other agents and combines the results. Anthropic uses this setup in its research system: one agent coordinates and several subagents search for information in parallel (Anthropic, 2025).
- Handoffs. One agent handles the conversation until it moves into another agent’s territory, then passes control along with the history. It’s the typical customer service pattern: a front-door agent that routes you to billing or returns. OpenAI’s Agents SDK implements it that way: a handoff shows up as just another tool the agent can use (OpenAI Agents SDK).
So that agents from different vendors can talk to each other, standards such as A2A have appeared. Google introduced it in April 2025 and later donated it to the Linux Foundation (Google, 2025).
When it pays off
The clearest case is tasks that can run in parallel and don’t fit in a single agent’s head. In its research system, Anthropic measured that the multi-agent version outperformed a single agent by 90.2% on its internal evaluation (Anthropic, 2025). Researching twenty companies at once is a good example: each subagent looks into one and the orchestrator compares.
It also helps when the parts really are different: different tools, different permissions, different risks. It makes sense for the agent that can issue a refund to be separate from the one that answers general questions, with fewer tools and stricter rules.
When it creates more problems than it solves
The same Anthropic post gives the price: its multi-agent systems used about 15 times more tokens than a regular chat. And it admits they’re a poor fit for tasks where every agent needs to share the same context or where agents depend heavily on each other.
Walden Yan, at Cognition, went further in a post bluntly titled Don’t Build Multi-Agents. His argument: every agent makes implicit decisions as it acts, and if it doesn’t share the full context with the others, those decisions clash (Yan, 2025). He recommends a single agent with the full context whenever possible.
The data partly backs him up. A team at the University of California, Berkeley analyzed more than 1,600 runs of seven popular multi-agent frameworks and sorted their failures into fourteen types, grouped into three families: system design issues, misalignment between agents, and lack of task verification (Cemri et al., 2025). Many of those failures wouldn’t exist with a single agent.
What changes on a phone call
In a voice conversation, the problems of multi-agent systems are more noticeable:
- Every handoff adds time. If the front-door agent has to think about where to route you and the next one has to read through the history, the person hears silence.
- The person needs to feel they’re still talking to the same service. If the voice, the tone or the way they’re addressed changes, it feels like being transferred from department to department.
- Context can’t get lost. Few things are more annoying than repeating your customer number after every handoff.
That’s why, in voice, we usually prefer a single agent facing the person, with one voice and one personality, that consults tools or specialized subagents behind the scenes when needed. The same agent always leads the conversation; what gets shared out is the work. It’s similar to the hybrid architecture we describe in three ways to build a voice agent.
How to decide
Before splitting one agent into several, ask yourself:
- Does the task have parts that can happen at the same time, or are they steps that depend on each other?
- Do the parts need different tools or permissions?
- How much context do they need to share?
- How much time and cost does each handoff add?
- Can you measure whether the split system works better than a single one?
If you don’t have a good answer to the last one, start with one agent. It’s easier to understand, test and fix. If it reaches its limit, your tests will tell you where to split it. That’s what we do when we design agents and when we evaluate them.
References