The idea in one minute#
Several agents are better than one in exactly three situations: the work splits into independent parts that can run in parallel, the work needs more context than one window can hold cleanly, or the parts need different permissions or owners. Outside those, a multi-agent system is a single agent with extra cost, extra latency and new ways to fail — agents misunderstanding each other is a bug class that one agent does not have.
When you do split, two things need design: the structure (who delegates to whom) and the interface (what exactly is passed, and what comes back).
A picture#
flowchart TB U[":i-user: <b>User</b>"] --> ORCH[":claude: <b>Orchestrator</b><br/><small>plans, delegates, synthesises</small>"] ORCH -->|"brief: goal, known facts,<br/>output format"| W1[":i-bot: <b>Worker: search</b><br/><small>fresh context, read-only tools</small>"] ORCH --> W2[":i-bot: <b>Worker: code</b><br/><small>sandbox, repo write</small>"] ORCH --> W3[":i-bot: <b>Worker: review</b><br/><small>different model, no write</small>"] W1 -->|"summary + sources"| ORCH W2 -->|"diff + test results"| ORCH W3 -->|"findings"| ORCH ORCH -->|"A2A task"| EXT[":a2a: <b>Partner's agent</b><br/><small>another organisation</small>"] EXT -->|"artifact"| ORCH W1 --> T1[":modelcontextprotocol: Search tools"] W2 --> T2[":i-box: Sandbox"] class U neutral class ORCH queue class W1,W2,W3 compute class EXT,T1 io class T2 warn
How it really works#
When to split, and when not to#
| Good reason | Example |
|---|---|
| Parallel, independent subtasks | Research ten companies; review twenty files |
| Context isolation | Exploring a large codebase without flooding the planner’s context |
| Separation of privilege | The agent that reads untrusted web pages has no write tools; the agent with write tools never sees raw web content |
| Independent verification | A reviewer with a fresh context and no stake in the first answer |
| Organisational boundary | Another team or vendor owns the agent and its data |
| Bad reason | Why |
|---|---|
| “Roles” mirroring a human org chart | Splits one coherent context into several ignorant ones |
| Tightly coupled steps | Each handoff loses information the next step needed |
| To seem sophisticated | Token use rises several-fold; so does latency |
A useful test: could you hand this subtask to a contractor with a one-page brief and get back what you need? If the brief would have to be the whole conversation, do not split.
Structures#
Orchestrator and workers. One agent plans, delegates subtasks to workers with fresh contexts, and synthesises their results. Workers do not talk to each other. This is the dependable default: one agent holds the whole picture and conflicts are resolved in one place.
Pipeline. Agents in a fixed sequence, each consuming the previous one’s output. This is a workflow whose steps happen to be agents; use it when the stages really are fixed.
Handoff. Control of the conversation passes from one agent to another — triage to billing to a human. One agent is active at a time, and the handoff must carry the state.
Generator and critic. One produces, another reviews against explicit criteria, and the loop repeats until the critic passes it or a budget ends. Most valuable when the critic has something objective to check — tests, a schema, a source to verify against.
Peer networks, where agents message each other freely, are hard to debug and to bound. Avoid them until you have a specific need that the structures above cannot meet.
The interface is the design#
Most multi-agent failures are briefing failures. A worker knows only what it is told, so a delegation message contains:
goal what outcome is wanted, and why it matters to the larger task
context what is already known; what has been tried and ruled out
scope what to do, and explicitly what not to do
constraints tools allowed, time and token budget, sources to prefer
output exact format; what counts as done; how to report uncertaintyAnd a result contains the answer, the evidence for it, and what was not verified. A worker that reports only success hides the gaps the orchestrator needs to know about. Treat a worker’s report as a claim to check, especially before acting on it.
Shared state#
Agents coordinate through shared artifacts more reliably than through conversation: a task board, files in a common workspace, a results table. Writes need ownership rules — two agents editing the same file is a merge conflict — so give each worker its own working area (a separate branch or directory) and merge deliberately.
Across organisational boundaries: A2A#
Inside your system, sub-agents are function calls with fresh contexts. Across a boundary you need a protocol, and that is A2A’s role: discover the other agent through its signed Agent Card, submit a task, receive status updates and artifacts, possibly over hours. What changes at the boundary is trust:
- You cannot see or control the other agent’s context, tools or model.
- Its output is untrusted input to yours, exactly like a web page.
- Its identity must be verified, and yours presented, on every call.
- Data you send it has left your control; apply the same rules as for any third-party API.
Cost and failure accounting#
single agent, 30 steps ≈ 1.1 M input tokens
orchestrator (10 steps) + 4 workers (15 steps each) ≈ 0.2 M + 4 × 0.4 M ≈ 1.8 M
wall-clock time: lower, if workers run in parallelExpect several times the tokens of a single agent, in exchange for wall-clock time and cleaner contexts. Budgets are set per worker and for the whole tree, so that a worker spawning workers cannot multiply without limit. Failures cascade the same way: decide whether one worker failing fails the task, is retried, or is reported as a gap in the final answer.
Observability#
One trace per task, spanning every agent. Each sub-agent is a child span carrying its brief, its result and its token use. Without this, “the system gave a wrong answer” cannot be traced to the agent and step that introduced the error.
Remember this#
- Split for parallelism, context isolation, privilege separation, independent review, or an organisational boundary. Otherwise do not.
- Orchestrator and workers is the default structure.
- The brief and the report format are the real design work.
- Another organisation’s agent is an untrusted input source and a data recipient.
- Budget the whole tree; trace the whole task.
Try it#
- Take a task you would give an agent. Apply the contractor test to three possible subtasks. Which pass?
- Write the brief for a “find all call sites of this function” worker. What must it return for the orchestrator to trust the result?
- Design a two-agent split in which the agent that reads untrusted content has no tool that can send data out. What passes between them?
Check yourself#
- What are the legitimate reasons to use more than one agent?
- Why are orchestrator-worker systems easier to debug than peer networks?
- How does trust differ between a sub-agent and an agent reached over A2A?