Pidoku

Anatomy of an Agent

Advanced 45 min Difficulty 3/5 Lesson 01 of 05

Prerequisites Tools and Protocols, Reliability Patterns

The idea in one minute#

An agent is a loop: the model reads the context, decides an action, the harness executes it, the result goes back into the context, repeat until the model says it is finished. The model supplies judgement. Everything else — which tools exist, what is in the context, what is allowed, when to stop, where state is saved — is the harness, and the harness is where engineering effort pays off. Two agents on the same model can differ enormously in capability because of it.

Before building one, check that you need one. If you can write the steps down in advance, you want a workflow, which is cheaper, faster and easier to test.

A picture#

flowchart LR
  GOAL[":i-user: <b>Goal</b>"] --> CTX
  subgraph HARNESS["The harness: your code"]
    direction LR
    CTX[":i-layers: <b>Build context</b><br/><small>instructions, history,<br/>tool list, memory</small>"] --> CALL[":i-brain: <b>Model call</b>"]
    CALL --> DEC{":i-split: <b>Reply is</b>"}
    DEC -->|"a tool call"| POL[":i-shield-check: <b>Policy check</b><br/><small>allowed? approval?<br/>budget left?</small>"]
    POL --> EXE[":i-wrench: <b>Execute</b><br/><small>tool, MCP, sandbox</small>"]
    EXE --> SAVE[":i-archive: <b>Checkpoint</b>"]
    SAVE --> CTX
  end
  DEC -->|"a final answer"| DONE[":i-check: <b>Result</b>"]
  POL -->|"denied or over budget"| STOP[":i-ban: <b>Stop, report</b>"]
  class GOAL,DONE neutral
  class CTX io
  class CALL compute
  class DEC,POL queue
  class EXE,STOP warn
  class SAVE memory

How it really works#

Workflow or agent#

WorkflowAgent
Control flowWritten by you, in codeChosen by the model at run time
StepsKnown in advanceDiscovered as it goes
Cost and latencyPredictableVariable, often 10–100× a single call
TestingStep by stepEnd state, statistically
Fails byA step failingWandering, looping, doing the wrong thing confidently
Good forClassify → retrieve → answer; extract → validate → storeDebug this failing test; research this question; resolve this ticket

Most useful systems are workflows with agentic steps: a fixed outer structure, with a model given freedom only where the path truly cannot be predicted. Common workflow shapes — prompt chaining, routing, parallel calls with voting, orchestrator and workers, a generator checked by an evaluator — cover a great deal before a free-running loop is needed.

The parts of the harness#

The system prompt states the role, the goal, the rules and how to use the tools. For long tasks it also says how to manage effort: when to plan, when to verify, when to stop and ask.

The tool set is the agent’s whole ability to act. A small set of general tools — read and write files, run a command, search, fetch a page — plus a few domain tools beats a hundred narrow ones. In 2026 the strongest general agents are built by giving the model a computer: a shell and a filesystem inside a sandbox, with domain capabilities added as MCP servers or as skills — folders of instructions and scripts the agent loads only when relevant.

The context builder assembles each call: instructions, tool definitions, relevant memory, the conversation and tool results so far. Its job over a long task is to keep the important things in and the bulk out — the subject of Context Engineering and Memory.

The policy layer sits between the model’s proposal and its execution. It validates arguments, checks the action against permissions, decides whether a human must approve, and enforces budgets. It is code, it cannot be argued with, and it is the main security control in an agentic system.

The executor runs the action: a function call, an MCP request, a command in a sandbox.

The checkpointer saves progress after every step so the task can resume after a crash, pause for an approval or be inspected afterwards.

The stop conditions. The model declaring completion is one of several. The others are the budgets of Reliability Patterns: steps, tokens, time, money, and a no-progress detector.

Verification is part of the loop#

An agent that cannot check its work drifts. The most reliable agents have a concrete signal to iterate against: tests that pass or fail, a compiler, a linter, a schema validator, a screenshot to compare. When designing an agent for a task, the first question is “how will it know it succeeded?” If the answer is “it will judge for itself”, add a check that does not depend on the model’s opinion.

Autonomy is a dial#

LevelThe humanTypical use
SuggestApproves every actionNew deployments; high-stakes domains
Act, ask for riskyApproves irreversible or external actionsMost production agents
Act within boundsReviews outcomes afterwardsSandboxed work with no external effects
Fully autonomousSets goals, auditsNarrow, well-evaluated, reversible tasks

Set the level per action class, not per agent: the same agent may read freely, write to its sandbox freely, and need approval to send an email or merge code. Ask for approval too often and people approve without reading, which is worse than not asking; reserve the prompt for actions that matter.

What an agent costs#

A 30-step coding task

context at step n   ≈ 12,000 (system + tools) + 1,500 × n (history grows)
total input tokens  ≈ Σ over 30 steps ≈ 1,060,000
total output tokens ≈ 30 × 600 = 18,000

without prefix caching   1.06 M × $3/M + 0.018 M × $15/M   ≈ $3.45
with 85% cache hits      (0.16 × $3 + 0.90 × $0.30) + $0.27 ≈ $1.02

Input is 98% of the tokens. This is why agent design is context design, and why the serving layer’s prefix cache matters more for agents than for chat.

Frameworks, briefly#

FrameworkIts model of an agent
LangGraphAn explicit graph of nodes and edges with built-in checkpointing and human-in-the-loop gates
Claude Agent SDKThe loop with a computer: shell, files, hooks, sub-agents, skills
OpenAI Agents SDKAgents that hand off to each other; sandbox execution and snapshots since its April 2026 revision
Google ADK, Microsoft Agent FrameworkTight integration with their clouds; A2A support
Your own loopFifty lines around a model API; often the right start

A framework saves the plumbing, not the design. Choose one by how it handles the hard parts: state, interruption, tracing, and how easily you can see exactly what was sent to the model.

Code#

The loop itself, with a scripted “model” so it runs offline. The policy check and the budget are the parts to study.

Go
// agent.go — the agent loop with a policy layer, a step budget and a no-progress detector.
package main

import (
	"fmt"
	"strings"
)

type action struct {
	tool, arg string
	final     string // set when the model is done
}

// model is scripted: a stand-in for an LLM call that looks at the last observation.
func model(history []string) action {
	last := history[len(history)-1]
	switch {
	case strings.HasPrefix(last, "goal:"):
		return action{tool: "read_file", arg: "report.csv"}
	case strings.Contains(last, "rows=3"):
		return action{tool: "send_email", arg: "[email protected]"}
	case strings.Contains(last, "DENIED"):
		return action{final: "Summary ready; emailing needs approval, so I stopped."}
	default:
		return action{tool: "read_file", arg: "report.csv"} // a loop, if nothing stops it
	}
}

var tools = map[string]struct {
	effect string // "read" or "external"
	run    func(arg string) string
}{
	"read_file":  {"read", func(arg string) string { return "read " + arg + ": rows=3" }},
	"send_email": {"external", func(arg string) string { return "sent to " + arg }},
}

// policy decides, in code, whether a proposed action may run.
func policy(a action, approved bool) (bool, string) {
	t, ok := tools[a.tool]
	if !ok {
		return false, "unknown tool"
	}
	if t.effect == "external" && !approved {
		return false, "external effect needs human approval"
	}
	return true, ""
}

func main() {
	const maxSteps = 8
	history := []string{"goal: summarise report.csv and send it to finance"}
	seen := map[string]int{}

	for step := 1; step <= maxSteps; step++ {
		a := model(history)
		if a.final != "" {
			fmt.Printf("step %d  done: %s\n", step, a.final)
			return
		}
		key := a.tool + "(" + a.arg + ")"
		if seen[key]++; seen[key] >= 3 {
			fmt.Printf("step %d  stopped: no progress, %s repeated\n", step, key)
			return
		}
		ok, why := policy(a, false)
		obs := "DENIED: " + why
		if ok {
			obs = tools[a.tool].run(a.arg)
		}
		fmt.Printf("step %d  %-34s → %s\n", step, key, obs)
		history = append(history, obs)
	}
	fmt.Println("stopped: step budget exhausted")
}

The model proposed sending an email; the policy refused; the model adapted. Nothing in the prompt was needed to make that happen.

Remember this#

  • An agent is model plus harness. The harness is where the engineering is.
  • Use a workflow when the steps are known; give the model freedom only where it is needed.
  • The policy layer is code between proposal and execution. It is the main control.
  • Give the agent a way to verify its own work.
  • Autonomy is set per action class.

Try it#

  1. Run agent.go. Change policy(a, false) to policy(a, true). What changes?
  2. Remove the DENIED case from model. Which mechanism stops the run now?
  3. For a task you would like to automate, list the tools, the verification signal and the action classes that need approval.

Check yourself#

  1. What distinguishes a workflow from an agent, and why does it matter for testing?
  2. Why is input the dominant cost of an agent?
  3. Why is asking for approval on every action a poor safety design?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom