Pidoku

State, Memory and Caching

Basic 45 min Difficulty 3/5 Lesson 03 of 06

Prerequisites Context and Retrieval

The idea in one minute#

The model is stateless, so the system holds all the state — and there are four different kinds that are often confused. Session state is the conversation. Task state is how far a long job has got. Memory is what should survive across sessions. Caches are copies kept to avoid paying twice. Each has a different lifetime, a different store and a different way of going wrong. Put each in the right place and model servers become interchangeable; mix them up and you get lost work, leaked data and bills that make no sense.

A picture#

flowchart TB
  APP[":i-bot: <b>Application or agent</b>"]
  subgraph STATE["State you own"]
    direction LR
    SS[(":redis: <b>Session</b><br/><small>messages, minutes to days</small>")]
    TS[(":temporal: <b>Task</b><br/><small>steps, checkpoints</small>")]
    MM[(":postgresql: <b>Memory</b><br/><small>facts, preferences, months</small>")]
  end
  subgraph CACHE["Caches: safe to lose"]
    direction LR
    C1[":i-zap: <b>Response cache</b><br/><small>exact match</small>"]
    C2[":i-search: <b>Semantic cache</b><br/><small>similar question</small>"]
    C3[":anthropic: <b>Prompt cache</b><br/><small>provider side, by prefix</small>"]
    C4[":vllm: <b>KV prefix cache</b><br/><small>on your GPUs</small>"]
  end
  APP --> SS
  APP --> TS
  APP --> MM
  APP --> C1 --> C2 --> C3
  C3 --- C4
  class APP compute
  class SS,TS,MM memory
  class C1,C2 queue
  class C3,C4 io

How it really works#

Four kinds of state#

KindHoldsLifetimeStoreFailure if mishandled
SessionMessages and tool results of one conversationMinutes to daysRedis or Postgres, keyed by sessionContext lost mid-conversation
TaskPlan, completed steps, pending tool callsUntil the task endsA durable workflow engine or a checkpoint tableAn hour of agent work repeated after a crash
MemoryFacts and preferences worth keepingWeeks to foreverPostgres rows, files, a vector indexStale or poisoned beliefs
CacheAnything recomputableUntil evictedIn-process, Redis, provider, GPUCost; with semantic caches, wrong answers

The rule from the design method applies: no state in the model process. Any replica must be able to serve the next call.

Session state and the growing prompt#

Every turn appends to the conversation, and the whole conversation is resent. Three techniques keep it bounded:

  • Trim tool results once they have been used; they are usually the bulkiest items.
  • Compact: when the conversation passes a threshold, have a model summarise the older part and continue with the summary plus the recent turns.
  • Offload: write large results to a file or store and keep only a reference in the context.

Compaction is lossy by design. Decide what must never be summarised away — the user’s goal, constraints, decisions already made — and keep those verbatim.

Memory#

Memory is a retrieval problem with a write path. Three decisions define it:

  • What to write. Explicit facts (“prefers metric units”), not whole transcripts.
  • Who can write. If text from a web page can become a memory, an attacker can plant an instruction that fires weeks later. Memory writes are a security boundary; see Agent Threats.
  • How it is read. Small memories are loaded into every prompt; large ones are searched. Many 2026 agents use plain files the agent reads and edits itself, because files are easy to inspect, diff and correct.

Scope every memory to its owner — user, team or organisation — and enforce that scope in the store, exactly as with retrieval permissions.

The four caches#

Response cache, exact match. Key on a hash of model, parameters and the full prompt. Safe, and only useful where identical requests recur: classification, extraction, batch jobs.

Semantic cache. Return a stored answer when a new question is similar to an old one. The savings are real and so are the risks:

warning

A semantic cache answers a question the user did not ask. “Cancel my order” and “can I cancel my order?” are close in vector space and need different responses. Never share a semantic cache across users or tenants, never use it for personalised or permissioned answers, and measure its false-hit rate before trusting it.

Prompt cache, provider side. Hosted APIs store the processed form of a prompt prefix and bill repeated prefixes at a fraction of the input price, with a lifetime of minutes. It is the largest and safest saving available, and it depends entirely on prompt layout: one changed character early in the prompt invalidates everything after it. Keep timestamps, user names and request IDs out of the prefix.

KV prefix cache, on your own GPUs. The same idea inside an engine: the attention state for a shared prefix is kept and reused, removing most of the prefill time. It only helps if the request reaches a replica that has that prefix, which is why inference-aware gateways route by prefix. See prefix caching and prefix-aware routing.

What the caches are worth#

An agent step: 40,000 input tokens, 600 output tokens, 30 steps per task

                     input price   input cost/step   task cost (30 steps)
no caching           $3.00 / M     $0.120            $3.87
85% prefix cached    $0.30 / M     $0.028            $1.11      ← same model, same answers

Caching the stable prefix cut the cost of the task by 70% and changed nothing about its behaviour. No other optimisation is that cheap.

What fills the box in 2026#

NeedOptions
Sessions and exact cachesRedis, Valkey, Postgres
Task stateTemporal, Restate, LangGraph checkpointers on Postgres
MemoryPostgres with pgvector, files in object storage, purpose-built memory services
Prefix cachingBuilt into hosted APIs and into vLLM, SGLang and NVIDIA Dynamo

Remember this#

  • Four kinds of state: session, task, memory, cache. Different lifetimes, different stores.
  • Nothing lives in the model process.
  • Compaction keeps sessions bounded; decide what must survive it verbatim.
  • Memory writes are a security boundary.
  • Prefix caching is the cheapest large saving. Semantic caching is the riskiest.

Try it#

  1. Recompute the cost table for 60 steps instead of 30, assuming the prompt grows by 1,500 tokens per step. How does the cached share change the picture?
  2. List what an assistant you use appears to remember across sessions. Where would you store each item, and who should be allowed to write it?
  3. Find one field in a prompt you maintain that would break prefix caching.

Check yourself#

  1. Why should task state not be kept in the same store, with the same lifetime, as session state?
  2. What makes a semantic cache dangerous in a multi-tenant system?
  3. Why does a timestamp at the top of a system prompt raise cost?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom