The idea in one minute#
A model can only produce text. It acts on the world through tool calling: you describe functions to it, it replies with a structured request to call one, your code runs the function and returns the result as more context. The model never executes anything. That separation is the most important fact in agent design, because it means every action passes through code you control.
Two open protocols standardise the wiring. MCP connects an agent to tools and data. A2A connects an agent to another agent. Both exist so that integrations are written once instead of once per framework.
A picture#
sequenceDiagram
participant U as User
participant H as Harness (your code)
participant M as Model
participant T as Tool (MCP server)
U->>H: "Refund order 8841"
H->>M: messages + tool definitions
M-->>H: tool_call: lookup_order(id=8841)
Note over H: validate arguments, check policy
H->>T: tools/call lookup_order
T-->>H: {status: delivered, total: 42.00}
H->>M: tool result
M-->>H: tool_call: issue_refund(id=8841, amount=42.00)
Note over H: needs approval: money moves
H->>U: Approve refund of 42.00?
U-->>H: yes
H->>T: tools/call issue_refund
T-->>H: {refund_id: r_19}
H->>M: tool result
M-->>H: "Refunded 42.00, reference r_19."
H->>U: final answerThe harness sits in the middle of every arrow. The model proposes; the harness decides.
How it really works#
Tool calling#
A tool definition has a name, a description and a JSON Schema for its arguments. The description is a prompt: the model chooses tools by reading it, so it should say when to use the tool, when not to, and what the result means.
name issue_refund
description Refund a delivered order. Use only after lookup_order confirms the status is
"delivered". Amount must not exceed the order total.
input { order_id: string, amount: number, reason: string }Design rules that hold in practice:
- Few, well-named tools. A model chooses better among 10 clear tools than 100 overlapping ones. Large catalogues are searched on demand rather than loaded whole.
- Validate in the harness. Arguments are model output, so they are untrusted input: check them against the schema and against business rules before executing.
- Make results useful to a reader. Return what the model needs to decide the next step, not a raw 5,000-line API response.
- Make errors instructive. “order_id must start with
ord_” lets the model correct itself. - Make writes idempotent. A retry must not refund twice. Pass an idempotency key.
- Classify by effect. Read-only, reversible write, irreversible write. The class decides whether a call needs approval.
Structured output is the same mechanism without the tool: constrain the reply to a JSON Schema so a program can consume it. Engines enforce the schema during decoding; your code still validates the values.
MCP: agent to tools#
Before MCP, each agent framework needed its own adapter for each system. MCP defines one client–server protocol so a server written once works with any client.
| Concept | What it is |
|---|---|
| Server | Exposes capabilities of one system: a database, a ticket tracker, a browser |
| Client | Lives in the agent’s host application and talks to servers |
| Tools | Functions the model may call |
| Resources | Data the application may read into context |
| Prompts | Reusable templates a user may invoke |
| Transports | stdio for a local process; Streamable HTTP for a remote service |
The 2026-07-28 specification is the one to design against. What changed matters architecturally:
- Stateless core. The initialize handshake and the session header are gone; every request carries its protocol version and capabilities. A remote MCP server is now an ordinary stateless HTTP service: load-balance it, autoscale it, deploy it without sticky sessions.
- Routable headers. The method and tool name travel in
Mcp-MethodandMcp-Nameheaders, so a gateway can authorise, rate-limit and meter per tool without parsing bodies. - Multi round-trip requests. When a tool needs input mid-call — a confirmation, a missing
field — the server returns
input_requiredand the client retries with the answer, instead of holding a stream open. - Cacheable lists. Tool and resource lists carry a time-to-live, so clients stop re-fetching them on every turn.
- Authorization tightened. OAuth 2.1 with issuer validation; Client ID Metadata Documents replace dynamic client registration; Enterprise-Managed Authorization lets an organisation’s identity provider decide which users may reach which servers.
- Extensions. Tasks (long-running calls you poll), MCP Apps (interactive UI) and enterprise authorization are official extensions rather than core.
Go, Python, TypeScript and C# have Tier 1 SDKs.
important
An MCP server’s tool descriptions and results are text that enters your model’s context. A server you did not write is a supplier you must vet, pin and monitor. See MCP and Tool Security.
A2A: agent to agent#
Sometimes the other side is not a function but another agent — a different team’s, a vendor’s — with its own model, tools and policies. A2A (v1.0, March 2026) standardises that conversation.
| Concept | What it is |
|---|---|
| Agent Card | A JSON document at a well-known URL describing an agent’s skills, endpoint and auth. In v1.0 it is cryptographically signed, so a caller can verify who published it |
| Task | A unit of work with a lifecycle: submitted, working, input-required, completed, failed |
| Message and artifact | What the agents exchange: text, files, structured data |
| Streaming and push | Progress updates for work that takes minutes or hours |
The distinction to keep: MCP gives your agent hands; A2A gives it colleagues. With MCP you see and control each call. With A2A you delegate an outcome and the other agent’s internals are opaque, so the trust question is different: you are trusting an organisation, not validating a function call.
Choosing how to connect something#
| Situation | Use |
|---|---|
| A function in your own codebase | A plain tool definition; no protocol needed |
| A system many agents or teams will use | An MCP server, behind your gateway |
| A local resource: files, a browser, a terminal | An MCP server over stdio, inside a sandbox |
| Another team’s or company’s agent | A2A |
| A bulk, deterministic operation | Code written by the agent and run in a sandbox, not thousands of tool calls |
That last row is a real 2026 pattern: rather than calling a tool a thousand times, the agent writes a short script that calls the API in a loop and runs it in a sandbox. It uses far fewer tokens and needs a place to execute safely — see Sandboxes and Tool Execution.
Remember this#
- The model proposes tool calls; your harness validates, authorises and executes them.
- Tool descriptions are prompts. Tool arguments and results are untrusted.
- MCP: agent to tools, stateless since 2026-07-28, routable by header.
- A2A: agent to agent, tasks with a lifecycle, signed Agent Cards.
- Classify tools by effect and gate the irreversible ones.
Try it#
- Write the tool definition for “send an email on the user’s behalf”. Which class of effect is it, and what must the harness check before running it?
- For three systems at your workplace, decide: plain tool, MCP server or A2A?
- Rewrite a vague error message from an API you use so that a model could fix its own call.
Check yourself#
- Why is it significant that the model never executes a tool itself?
- What does a stateless MCP core change about deploying a remote server?
- When would you choose A2A over MCP?