Pidoku

Multi-Agent Systems

Advanced 45 min Difficulty 4/5 Lesson 04 of 05

Prerequisites Context Engineering and Memory, Tools and Protocols

The idea in one minute#

Several agents are better than one in exactly three situations: the work splits into independent parts that can run in parallel, the work needs more context than one window can hold cleanly, or the parts need different permissions or owners. Outside those, a multi-agent system is a single agent with extra cost, extra latency and new ways to fail — agents misunderstanding each other is a bug class that one agent does not have.

When you do split, two things need design: the structure (who delegates to whom) and the interface (what exactly is passed, and what comes back).

A picture#

flowchart TB
  U[":i-user: <b>User</b>"] --> ORCH[":claude: <b>Orchestrator</b><br/><small>plans, delegates, synthesises</small>"]
  ORCH -->|"brief: goal, known facts,<br/>output format"| W1[":i-bot: <b>Worker: search</b><br/><small>fresh context, read-only tools</small>"]
  ORCH --> W2[":i-bot: <b>Worker: code</b><br/><small>sandbox, repo write</small>"]
  ORCH --> W3[":i-bot: <b>Worker: review</b><br/><small>different model, no write</small>"]
  W1 -->|"summary + sources"| ORCH
  W2 -->|"diff + test results"| ORCH
  W3 -->|"findings"| ORCH
  ORCH -->|"A2A task"| EXT[":a2a: <b>Partner's agent</b><br/><small>another organisation</small>"]
  EXT -->|"artifact"| ORCH
  W1 --> T1[":modelcontextprotocol: Search tools"]
  W2 --> T2[":i-box: Sandbox"]
  class U neutral
  class ORCH queue
  class W1,W2,W3 compute
  class EXT,T1 io
  class T2 warn

How it really works#

When to split, and when not to#

Good reasonExample
Parallel, independent subtasksResearch ten companies; review twenty files
Context isolationExploring a large codebase without flooding the planner’s context
Separation of privilegeThe agent that reads untrusted web pages has no write tools; the agent with write tools never sees raw web content
Independent verificationA reviewer with a fresh context and no stake in the first answer
Organisational boundaryAnother team or vendor owns the agent and its data
Bad reasonWhy
“Roles” mirroring a human org chartSplits one coherent context into several ignorant ones
Tightly coupled stepsEach handoff loses information the next step needed
To seem sophisticatedToken use rises several-fold; so does latency

A useful test: could you hand this subtask to a contractor with a one-page brief and get back what you need? If the brief would have to be the whole conversation, do not split.

Structures#

Orchestrator and workers. One agent plans, delegates subtasks to workers with fresh contexts, and synthesises their results. Workers do not talk to each other. This is the dependable default: one agent holds the whole picture and conflicts are resolved in one place.

Pipeline. Agents in a fixed sequence, each consuming the previous one’s output. This is a workflow whose steps happen to be agents; use it when the stages really are fixed.

Handoff. Control of the conversation passes from one agent to another — triage to billing to a human. One agent is active at a time, and the handoff must carry the state.

Generator and critic. One produces, another reviews against explicit criteria, and the loop repeats until the critic passes it or a budget ends. Most valuable when the critic has something objective to check — tests, a schema, a source to verify against.

Peer networks, where agents message each other freely, are hard to debug and to bound. Avoid them until you have a specific need that the structures above cannot meet.

The interface is the design#

Most multi-agent failures are briefing failures. A worker knows only what it is told, so a delegation message contains:

goal              what outcome is wanted, and why it matters to the larger task
context           what is already known; what has been tried and ruled out
scope             what to do, and explicitly what not to do
constraints       tools allowed, time and token budget, sources to prefer
output            exact format; what counts as done; how to report uncertainty

And a result contains the answer, the evidence for it, and what was not verified. A worker that reports only success hides the gaps the orchestrator needs to know about. Treat a worker’s report as a claim to check, especially before acting on it.

Shared state#

Agents coordinate through shared artifacts more reliably than through conversation: a task board, files in a common workspace, a results table. Writes need ownership rules — two agents editing the same file is a merge conflict — so give each worker its own working area (a separate branch or directory) and merge deliberately.

Across organisational boundaries: A2A#

Inside your system, sub-agents are function calls with fresh contexts. Across a boundary you need a protocol, and that is A2A’s role: discover the other agent through its signed Agent Card, submit a task, receive status updates and artifacts, possibly over hours. What changes at the boundary is trust:

  • You cannot see or control the other agent’s context, tools or model.
  • Its output is untrusted input to yours, exactly like a web page.
  • Its identity must be verified, and yours presented, on every call.
  • Data you send it has left your control; apply the same rules as for any third-party API.

Cost and failure accounting#

single agent, 30 steps                               ≈ 1.1 M input tokens
orchestrator (10 steps) + 4 workers (15 steps each)  ≈ 0.2 M + 4 × 0.4 M ≈ 1.8 M
                                                       wall-clock time: lower, if workers run in parallel

Expect several times the tokens of a single agent, in exchange for wall-clock time and cleaner contexts. Budgets are set per worker and for the whole tree, so that a worker spawning workers cannot multiply without limit. Failures cascade the same way: decide whether one worker failing fails the task, is retried, or is reported as a gap in the final answer.

Observability#

One trace per task, spanning every agent. Each sub-agent is a child span carrying its brief, its result and its token use. Without this, “the system gave a wrong answer” cannot be traced to the agent and step that introduced the error.

Remember this#

  • Split for parallelism, context isolation, privilege separation, independent review, or an organisational boundary. Otherwise do not.
  • Orchestrator and workers is the default structure.
  • The brief and the report format are the real design work.
  • Another organisation’s agent is an untrusted input source and a data recipient.
  • Budget the whole tree; trace the whole task.

Try it#

  1. Take a task you would give an agent. Apply the contractor test to three possible subtasks. Which pass?
  2. Write the brief for a “find all call sites of this function” worker. What must it return for the orchestrator to trust the result?
  3. Design a two-agent split in which the agent that reads untrusted content has no tool that can send data out. What passes between them?

Check yourself#

  1. What are the legitimate reasons to use more than one agent?
  2. Why are orchestrator-worker systems easier to debug than peer networks?
  3. How does trust differ between a sub-agent and an agent reached over A2A?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom