The idea in one minute#
An agent’s context grows with every step, and three things go wrong as it does: the window fills, the bill grows quadratically, and the model’s attention degrades — it starts to lose track of instructions and earlier findings long before the window is technically full. Context engineering is keeping the context small and sharp for the whole task. There are four moves: select what enters, compress what has been used, isolate bulky work in sub-agents, and persist what must outlive the window in external memory.
A picture#
flowchart TB
subgraph WIN["The context window, every call"]
direction LR
A["Stable prefix<br/><small>system prompt, tools</small>"] --- B["Memory<br/><small>loaded notes</small>"] --- C["Compacted history<br/><small>summary of old steps</small>"] --- D["Recent steps<br/><small>verbatim</small>"]
end
SEL[":i-funnel: <b>Select</b><br/><small>load tools and docs on demand</small>"] --> WIN
WIN --> CMP[":i-recycle: <b>Compress</b><br/><small>summarise, clear old tool results</small>"]
CMP --> WIN
WIN --> ISO[":i-bot: <b>Isolate</b><br/><small>sub-agent does bulky work,<br/>returns a summary</small>"]
ISO --> WIN
WIN --> PER[(":i-file-text: <b>Persist</b><br/><small>notes, plans, files</small>")]
PER --> WIN
class A,B memory
class C,D neutral
class SEL,CMP queue
class ISO compute
class PER memoryHow it really works#
Why long contexts degrade#
Three separate effects, often confused:
- Capacity: the window has a hard limit.
- Cost and latency: every token is processed on every step; a 150,000-token context makes each step slow and expensive even when nothing in it is needed.
- Attention: recall of details buried in the middle of a long context is measurably worse than at the ends, and contradictory or stale content actively misleads. This sets in well below the hard limit.
So “the model has a million-token window” is not a design. The target is the smallest context that contains what this step needs.
Select: let things in on demand#
- Tools on demand. Do not load 200 tool definitions; give the agent a search tool over the catalogue and load the few it asks for.
- Skills. Package procedural knowledge as folders — a short description always in context, the full instructions and scripts read only when the task calls for them.
- Just-in-time retrieval. Keep references (file paths, URLs, query handles) in context and let the agent fetch contents when needed, rather than pasting everything up front.
- Bounded tool results. Tools return the first page, a summary, or a path to the full output — never an unbounded dump.
Compress: shrink what has been used#
- Clear stale tool results. Once a file has been read and acted on, its 8,000 tokens can be replaced by a one-line stub. This is cheap, safe and usually the largest saving.
- Compaction. When the context passes a threshold, a model call summarises the older steps — goal, decisions made, facts established, work remaining — and the task continues from the summary plus the most recent steps.
Compaction loses information, and a second compaction summarises a summary. Protect what must not degrade: keep the original goal, hard constraints and user corrections verbatim outside the summarised region, and have the agent write important findings to a file before they can be summarised away.
There is a cost interaction: compaction rewrites the prefix, so the prompt cache is lost at that moment. Compact at a few well-chosen points, not continuously.
Isolate: sub-agents as context firewalls#
A sub-agent is given a narrow task and a fresh context, does the bulky work — reads forty files, runs a dozen searches — and returns a short result. The parent’s context receives 1,000 tokens instead of 80,000. The primary benefit is not parallelism; it is that the mess stays out of the main context.
Cost to respect: the sub-agent starts with no knowledge, so the parent must brief it fully — goal, what is already known, what to return, in what form. A vague brief gets a vague result, and the sub-agent’s total token use is still paid for. Multi-Agent Systems continues this.
Persist: memory outside the window#
| Memory | Holds | Typical form |
|---|---|---|
| Working notes | Plan, to-do list, progress on this task | A file the agent rewrites as it goes |
| Project memory | Conventions, commands, architecture notes | Instruction files loaded at start (the AGENTS.md convention) |
| User memory | Preferences, standing corrections | Small records, loaded into every session |
| Episodic memory | What happened in past tasks | Searchable store of summaries |
| Knowledge | Documents, code | The retrieval layer |
A to-do list the agent maintains is the simplest and most effective of these: it survives compaction, keeps a long task oriented, and lets a resumed task know where it was. Files work well as memory because they are inspectable, diffable and correctable by a human.
Three rules for memory design:
- Write deliberately. Store conclusions and corrections, not transcripts.
- Scope and expire. Each memory has an owner and a lifetime; stale memory is worse than none.
- Guard the write path. Whatever can write memory can influence every future session. If content from a web page or an email can cause a memory write, that is a persistent injection channel — see Agent Threats.
Layout for the cache#
The four moves must respect the prefix cache, or you save tokens and lose money:
position content changes cached?
──────── ─────── ─────── ───────
first system prompt, tool definitions per release yes, long-lived
then loaded memory, project notes per session yes, for the session
then compacted summary at each compaction yes, between compactions
last recent steps, new tool results every step the tail is always newAppend, do not edit. Changing something early in the context — reordering tools, updating a timestamp, editing an old message — invalidates the cache for everything after it.
How to tell it is working#
Track per task: peak context size, tokens per step, cache hit rate, number of compactions, and success rate against task length. If success falls as tasks get longer, context is the first suspect: read a failing trace at the step where it went wrong and look at what the model was actually shown.
Remember this#
- Aim for the smallest context that contains what the step needs.
- Four moves: select, compress, isolate, persist.
- Clearing old tool results is the cheapest compression; compaction is lossy, so protect the goal and constraints.
- Sub-agents keep bulky work out of the main context; brief them fully.
- Memory writes are deliberate, scoped and guarded.
- Append-only context keeps the prefix cache alive.
Try it#
- Take a long agent transcript and mark which tokens were still needed at the final step. What share of the context was dead weight?
- Write the compaction prompt for a coding agent. Which five items must the summary contain?
- Design the memory for a support agent: what is written, by whom, with what lifetime?
Check yourself#
- Name the three distinct problems of a long context.
- Why does continuous compaction cost more than occasional compaction?
- What is the main benefit of a sub-agent, if not speed?