Pidoku

Tools and Protocols

Basic 50 min Difficulty 3/5 Lesson 04 of 06

Prerequisites What a Model Changes

The idea in one minute#

A model can only produce text. It acts on the world through tool calling: you describe functions to it, it replies with a structured request to call one, your code runs the function and returns the result as more context. The model never executes anything. That separation is the most important fact in agent design, because it means every action passes through code you control.

Two open protocols standardise the wiring. MCP connects an agent to tools and data. A2A connects an agent to another agent. Both exist so that integrations are written once instead of once per framework.

A picture#

sequenceDiagram
  participant U as User
  participant H as Harness (your code)
  participant M as Model
  participant T as Tool (MCP server)
  U->>H: "Refund order 8841"
  H->>M: messages + tool definitions
  M-->>H: tool_call: lookup_order(id=8841)
  Note over H: validate arguments, check policy
  H->>T: tools/call lookup_order
  T-->>H: {status: delivered, total: 42.00}
  H->>M: tool result
  M-->>H: tool_call: issue_refund(id=8841, amount=42.00)
  Note over H: needs approval: money moves
  H->>U: Approve refund of 42.00?
  U-->>H: yes
  H->>T: tools/call issue_refund
  T-->>H: {refund_id: r_19}
  H->>M: tool result
  M-->>H: "Refunded 42.00, reference r_19."
  H->>U: final answer

The harness sits in the middle of every arrow. The model proposes; the harness decides.

How it really works#

Tool calling#

A tool definition has a name, a description and a JSON Schema for its arguments. The description is a prompt: the model chooses tools by reading it, so it should say when to use the tool, when not to, and what the result means.

name         issue_refund
description  Refund a delivered order. Use only after lookup_order confirms the status is
             "delivered". Amount must not exceed the order total.
input        { order_id: string, amount: number, reason: string }

Design rules that hold in practice:

  • Few, well-named tools. A model chooses better among 10 clear tools than 100 overlapping ones. Large catalogues are searched on demand rather than loaded whole.
  • Validate in the harness. Arguments are model output, so they are untrusted input: check them against the schema and against business rules before executing.
  • Make results useful to a reader. Return what the model needs to decide the next step, not a raw 5,000-line API response.
  • Make errors instructive. “order_id must start with ord_” lets the model correct itself.
  • Make writes idempotent. A retry must not refund twice. Pass an idempotency key.
  • Classify by effect. Read-only, reversible write, irreversible write. The class decides whether a call needs approval.

Structured output is the same mechanism without the tool: constrain the reply to a JSON Schema so a program can consume it. Engines enforce the schema during decoding; your code still validates the values.

MCP: agent to tools#

Before MCP, each agent framework needed its own adapter for each system. MCP defines one client–server protocol so a server written once works with any client.

ConceptWhat it is
ServerExposes capabilities of one system: a database, a ticket tracker, a browser
ClientLives in the agent’s host application and talks to servers
ToolsFunctions the model may call
ResourcesData the application may read into context
PromptsReusable templates a user may invoke
Transportsstdio for a local process; Streamable HTTP for a remote service

The 2026-07-28 specification is the one to design against. What changed matters architecturally:

  • Stateless core. The initialize handshake and the session header are gone; every request carries its protocol version and capabilities. A remote MCP server is now an ordinary stateless HTTP service: load-balance it, autoscale it, deploy it without sticky sessions.
  • Routable headers. The method and tool name travel in Mcp-Method and Mcp-Name headers, so a gateway can authorise, rate-limit and meter per tool without parsing bodies.
  • Multi round-trip requests. When a tool needs input mid-call — a confirmation, a missing field — the server returns input_required and the client retries with the answer, instead of holding a stream open.
  • Cacheable lists. Tool and resource lists carry a time-to-live, so clients stop re-fetching them on every turn.
  • Authorization tightened. OAuth 2.1 with issuer validation; Client ID Metadata Documents replace dynamic client registration; Enterprise-Managed Authorization lets an organisation’s identity provider decide which users may reach which servers.
  • Extensions. Tasks (long-running calls you poll), MCP Apps (interactive UI) and enterprise authorization are official extensions rather than core.

Go, Python, TypeScript and C# have Tier 1 SDKs.

important

An MCP server’s tool descriptions and results are text that enters your model’s context. A server you did not write is a supplier you must vet, pin and monitor. See MCP and Tool Security.

A2A: agent to agent#

Sometimes the other side is not a function but another agent — a different team’s, a vendor’s — with its own model, tools and policies. A2A (v1.0, March 2026) standardises that conversation.

ConceptWhat it is
Agent CardA JSON document at a well-known URL describing an agent’s skills, endpoint and auth. In v1.0 it is cryptographically signed, so a caller can verify who published it
TaskA unit of work with a lifecycle: submitted, working, input-required, completed, failed
Message and artifactWhat the agents exchange: text, files, structured data
Streaming and pushProgress updates for work that takes minutes or hours

The distinction to keep: MCP gives your agent hands; A2A gives it colleagues. With MCP you see and control each call. With A2A you delegate an outcome and the other agent’s internals are opaque, so the trust question is different: you are trusting an organisation, not validating a function call.

Choosing how to connect something#

SituationUse
A function in your own codebaseA plain tool definition; no protocol needed
A system many agents or teams will useAn MCP server, behind your gateway
A local resource: files, a browser, a terminalAn MCP server over stdio, inside a sandbox
Another team’s or company’s agentA2A
A bulk, deterministic operationCode written by the agent and run in a sandbox, not thousands of tool calls

That last row is a real 2026 pattern: rather than calling a tool a thousand times, the agent writes a short script that calls the API in a loop and runs it in a sandbox. It uses far fewer tokens and needs a place to execute safely — see Sandboxes and Tool Execution.

Remember this#

  • The model proposes tool calls; your harness validates, authorises and executes them.
  • Tool descriptions are prompts. Tool arguments and results are untrusted.
  • MCP: agent to tools, stateless since 2026-07-28, routable by header.
  • A2A: agent to agent, tasks with a lifecycle, signed Agent Cards.
  • Classify tools by effect and gate the irreversible ones.

Try it#

  1. Write the tool definition for “send an email on the user’s behalf”. Which class of effect is it, and what must the harness check before running it?
  2. For three systems at your workplace, decide: plain tool, MCP server or A2A?
  3. Rewrite a vague error message from an API you use so that a model could fix its own call.

Check yourself#

  1. Why is it significant that the model never executes a tool itself?
  2. What does a stateless MCP core change about deploying a remote server?
  3. When would you choose A2A over MCP?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom