The idea in one minute#
An agent is a loop: the model reads the context, decides an action, the harness executes it, the result goes back into the context, repeat until the model says it is finished. The model supplies judgement. Everything else — which tools exist, what is in the context, what is allowed, when to stop, where state is saved — is the harness, and the harness is where engineering effort pays off. Two agents on the same model can differ enormously in capability because of it.
Before building one, check that you need one. If you can write the steps down in advance, you want a workflow, which is cheaper, faster and easier to test.
A picture#
flowchart LR
GOAL[":i-user: <b>Goal</b>"] --> CTX
subgraph HARNESS["The harness: your code"]
direction LR
CTX[":i-layers: <b>Build context</b><br/><small>instructions, history,<br/>tool list, memory</small>"] --> CALL[":i-brain: <b>Model call</b>"]
CALL --> DEC{":i-split: <b>Reply is</b>"}
DEC -->|"a tool call"| POL[":i-shield-check: <b>Policy check</b><br/><small>allowed? approval?<br/>budget left?</small>"]
POL --> EXE[":i-wrench: <b>Execute</b><br/><small>tool, MCP, sandbox</small>"]
EXE --> SAVE[":i-archive: <b>Checkpoint</b>"]
SAVE --> CTX
end
DEC -->|"a final answer"| DONE[":i-check: <b>Result</b>"]
POL -->|"denied or over budget"| STOP[":i-ban: <b>Stop, report</b>"]
class GOAL,DONE neutral
class CTX io
class CALL compute
class DEC,POL queue
class EXE,STOP warn
class SAVE memoryHow it really works#
Workflow or agent#
| Workflow | Agent | |
|---|---|---|
| Control flow | Written by you, in code | Chosen by the model at run time |
| Steps | Known in advance | Discovered as it goes |
| Cost and latency | Predictable | Variable, often 10–100× a single call |
| Testing | Step by step | End state, statistically |
| Fails by | A step failing | Wandering, looping, doing the wrong thing confidently |
| Good for | Classify → retrieve → answer; extract → validate → store | Debug this failing test; research this question; resolve this ticket |
Most useful systems are workflows with agentic steps: a fixed outer structure, with a model given freedom only where the path truly cannot be predicted. Common workflow shapes — prompt chaining, routing, parallel calls with voting, orchestrator and workers, a generator checked by an evaluator — cover a great deal before a free-running loop is needed.
The parts of the harness#
The system prompt states the role, the goal, the rules and how to use the tools. For long tasks it also says how to manage effort: when to plan, when to verify, when to stop and ask.
The tool set is the agent’s whole ability to act. A small set of general tools — read and write files, run a command, search, fetch a page — plus a few domain tools beats a hundred narrow ones. In 2026 the strongest general agents are built by giving the model a computer: a shell and a filesystem inside a sandbox, with domain capabilities added as MCP servers or as skills — folders of instructions and scripts the agent loads only when relevant.
The context builder assembles each call: instructions, tool definitions, relevant memory, the conversation and tool results so far. Its job over a long task is to keep the important things in and the bulk out — the subject of Context Engineering and Memory.
The policy layer sits between the model’s proposal and its execution. It validates arguments, checks the action against permissions, decides whether a human must approve, and enforces budgets. It is code, it cannot be argued with, and it is the main security control in an agentic system.
The executor runs the action: a function call, an MCP request, a command in a sandbox.
The checkpointer saves progress after every step so the task can resume after a crash, pause for an approval or be inspected afterwards.
The stop conditions. The model declaring completion is one of several. The others are the budgets of Reliability Patterns: steps, tokens, time, money, and a no-progress detector.
Verification is part of the loop#
An agent that cannot check its work drifts. The most reliable agents have a concrete signal to iterate against: tests that pass or fail, a compiler, a linter, a schema validator, a screenshot to compare. When designing an agent for a task, the first question is “how will it know it succeeded?” If the answer is “it will judge for itself”, add a check that does not depend on the model’s opinion.
Autonomy is a dial#
| Level | The human | Typical use |
|---|---|---|
| Suggest | Approves every action | New deployments; high-stakes domains |
| Act, ask for risky | Approves irreversible or external actions | Most production agents |
| Act within bounds | Reviews outcomes afterwards | Sandboxed work with no external effects |
| Fully autonomous | Sets goals, audits | Narrow, well-evaluated, reversible tasks |
Set the level per action class, not per agent: the same agent may read freely, write to its sandbox freely, and need approval to send an email or merge code. Ask for approval too often and people approve without reading, which is worse than not asking; reserve the prompt for actions that matter.
What an agent costs#
A 30-step coding task
context at step n ≈ 12,000 (system + tools) + 1,500 × n (history grows)
total input tokens ≈ Σ over 30 steps ≈ 1,060,000
total output tokens ≈ 30 × 600 = 18,000
without prefix caching 1.06 M × $3/M + 0.018 M × $15/M ≈ $3.45
with 85% cache hits (0.16 × $3 + 0.90 × $0.30) + $0.27 ≈ $1.02Input is 98% of the tokens. This is why agent design is context design, and why the serving layer’s prefix cache matters more for agents than for chat.
Frameworks, briefly#
| Framework | Its model of an agent |
|---|---|
| LangGraph | An explicit graph of nodes and edges with built-in checkpointing and human-in-the-loop gates |
| Claude Agent SDK | The loop with a computer: shell, files, hooks, sub-agents, skills |
| OpenAI Agents SDK | Agents that hand off to each other; sandbox execution and snapshots since its April 2026 revision |
| Google ADK, Microsoft Agent Framework | Tight integration with their clouds; A2A support |
| Your own loop | Fifty lines around a model API; often the right start |
A framework saves the plumbing, not the design. Choose one by how it handles the hard parts: state, interruption, tracing, and how easily you can see exactly what was sent to the model.
Code#
The loop itself, with a scripted “model” so it runs offline. The policy check and the budget are the parts to study.
// agent.go — the agent loop with a policy layer, a step budget and a no-progress detector.
package main
import (
"fmt"
"strings"
)
type action struct {
tool, arg string
final string // set when the model is done
}
// model is scripted: a stand-in for an LLM call that looks at the last observation.
func model(history []string) action {
last := history[len(history)-1]
switch {
case strings.HasPrefix(last, "goal:"):
return action{tool: "read_file", arg: "report.csv"}
case strings.Contains(last, "rows=3"):
return action{tool: "send_email", arg: "[email protected]"}
case strings.Contains(last, "DENIED"):
return action{final: "Summary ready; emailing needs approval, so I stopped."}
default:
return action{tool: "read_file", arg: "report.csv"} // a loop, if nothing stops it
}
}
var tools = map[string]struct {
effect string // "read" or "external"
run func(arg string) string
}{
"read_file": {"read", func(arg string) string { return "read " + arg + ": rows=3" }},
"send_email": {"external", func(arg string) string { return "sent to " + arg }},
}
// policy decides, in code, whether a proposed action may run.
func policy(a action, approved bool) (bool, string) {
t, ok := tools[a.tool]
if !ok {
return false, "unknown tool"
}
if t.effect == "external" && !approved {
return false, "external effect needs human approval"
}
return true, ""
}
func main() {
const maxSteps = 8
history := []string{"goal: summarise report.csv and send it to finance"}
seen := map[string]int{}
for step := 1; step <= maxSteps; step++ {
a := model(history)
if a.final != "" {
fmt.Printf("step %d done: %s\n", step, a.final)
return
}
key := a.tool + "(" + a.arg + ")"
if seen[key]++; seen[key] >= 3 {
fmt.Printf("step %d stopped: no progress, %s repeated\n", step, key)
return
}
ok, why := policy(a, false)
obs := "DENIED: " + why
if ok {
obs = tools[a.tool].run(a.arg)
}
fmt.Printf("step %d %-34s → %s\n", step, key, obs)
history = append(history, obs)
}
fmt.Println("stopped: step budget exhausted")
}The model proposed sending an email; the policy refused; the model adapted. Nothing in the prompt was needed to make that happen.
Remember this#
- An agent is model plus harness. The harness is where the engineering is.
- Use a workflow when the steps are known; give the model freedom only where it is needed.
- The policy layer is code between proposal and execution. It is the main control.
- Give the agent a way to verify its own work.
- Autonomy is set per action class.
Try it#
- Run
agent.go. Changepolicy(a, false)topolicy(a, true). What changes? - Remove the
DENIEDcase frommodel. Which mechanism stops the run now? - For a task you would like to automate, list the tools, the verification signal and the action classes that need approval.
Check yourself#
- What distinguishes a workflow from an agent, and why does it matter for testing?
- Why is input the dominant cost of an agent?
- Why is asking for approval on every action a poor safety design?