The idea in one minute#
Put a language model in the request path and five assumptions of ordinary system design break. The component is slow (seconds, not milliseconds), metered by size (tokens, not requests), probabilistic (the same input can give a different output), stateless but context-hungry (it remembers nothing, so you resend everything) and steerable by its input (text it reads can change what it does). An AI system design is an ordinary design plus an explicit answer to each of those five.
A picture#
flowchart LR
subgraph CLASSIC["An ordinary request"]
C1[":i-user: Client"] --> C2[":nginx: API"] --> C3[(":postgresql: Database")]
end
subgraph AI["The same request with a model in it"]
A1[":i-user: Client"] --> A2[":envoyproxy: <b>Gateway</b><br/><small>counts tokens, not requests</small>"]
A2 --> A3[":i-layers: <b>Context builder</b><br/><small>history, documents, tools</small>"]
A3 --> A4[":i-brain: <b>Model</b><br/><small>1-30 s, streamed, non-deterministic</small>"]
A4 --> A5[":i-shield-check: <b>Output check</b><br/><small>parse, validate, filter</small>"]
A4 -.->|"may ask to act"| A6[":i-wrench: Tools"]
A6 -.-> A3
end
class C1,A1 neutral
class C2,A2 queue
class C3 memory
class A3 io
class A4 compute
class A5,A6 warnThree boxes became six, and one arrow now points backwards. Each new box exists because of one of the five differences.
How it really works#
1. Slow: latency is a shape, not a number#
A database query returns in milliseconds, all at once. A model returns a stream. Two numbers describe it:
- TTFT (time to first token) — how long the user stares at nothing. It grows with the length of the input.
- TPOT (time per output token) — how fast text appears afterwards. It decides how long a long answer takes.
total ≈ TTFT + output_tokens × TPOT
a chat reply 0.4 s + 300 × 0.02 s ≈ 6.4 s
a tool decision 0.4 s + 40 × 0.02 s ≈ 1.2 s
an agent task 30 model calls × ~3 s ≈ 90 s, before any tool runsConsequences for the design: everything streams; timeouts are tens of seconds; connection slots are held for a long time, so concurrency limits matter more than request rate; and a 30-step agent is a background job, not a request. Latency and throughput covers the measurements in depth.
2. Metered by size: the token is the unit#
Cost, latency, rate limits and capacity are all proportional to tokens, and one request can be 50 tokens or 500,000. A design that counts requests will be wrong by four orders of magnitude in both directions.
| Ordinary design counts | An AI design counts |
|---|---|
| Requests per second | Input tokens/s and output tokens/s, separately |
| Rate limit: 100 requests/min | Rate limit: tokens per minute, plus a spend budget |
| Payload size, rarely | Context length, on every call |
| Cost per request ≈ constant | Cost per request varies 10,000× |
Output tokens cost several times more than input tokens, on hosted APIs and on your own GPUs alike, because output is generated one token at a time while input is processed in parallel. Input that repeats — a long system prompt, a document — can be cached and billed at a fraction. That one fact shapes prompt layout: stable content first, changing content last.
3. Probabilistic: correctness becomes a rate#
The same prompt can produce different text, and “correct” is often a judgement, not an equality check. So:
- Tests become evaluations. You measure a pass rate on a dataset and gate releases on it not dropping. See Evaluation and Quality.
- Outputs are validated, not trusted. Anything a program will consume is parsed against a schema, and the design says what happens when parsing fails.
- Retries are a tool and a cost. Retrying a failed parse usually works; it also doubles the bill for that request.
- A model upgrade is a behaviour change, even with identical code. Models are versioned and rolled out like code, with the evaluation as the gate.
4. Stateless but context-hungry: you carry the memory#
The model keeps nothing between calls. Everything it should know — instructions, conversation, documents, tool results — is assembled into the context window on every call. That makes the context builder the real heart of the application: it decides what the model sees, and the model can only be as good as that. The window is finite and every token in it costs money and time, so “what goes in the context” is a design decision with a budget, covered in Context and Retrieval.
5. Steerable by its input: the model is not a trusted component#
A model does not separate instructions from data. Text arriving in a retrieved document, a web page or a tool result can redirect it — prompt injection — and no filter reliably stops that. The architectural consequence is simple to state and easy to forget:
important
Treat model output as untrusted input. The permissions a model’s output can exercise must be limited by the system around it, never by the instructions inside it.
This is why the picture has an output check, why tools sit behind their own authorization, and why Security by Design is part of this course and not an appendix. The whole of AI Security Engineering follows from this one property.
What does not change#
Load balancing, queues, idempotency, caching, replication, observability, least privilege. Every one of them is still needed, and most AI outages are ordinary outages: a dependency timed out, a queue filled, a deploy was not rolled back. The model adds to classic design; it does not replace it.
The three shapes of AI system#
Almost every product is one of three shapes, in rising order of difficulty:
| Shape | The model | Typical example | Hard part |
|---|---|---|---|
| Single call | Transforms input to output once | Classify, extract, summarise | Cost and quality at volume |
| Retrieval-augmented | Answers using context you fetch | Assistant over company documents | What goes in the context |
| Agentic | Chooses actions in a loop until done | Coding agent, support agent | State, safety, long-running work |
Design for the simplest shape that solves the problem. Each step up multiplies cost, latency and attack surface.
Remember this#
- Five differences: slow, token-metered, probabilistic, context-hungry, steerable by input.
- Count tokens, split by input and output. Never requests.
- Correctness is a measured rate, guarded by evaluations.
- The context builder decides what the model knows; the system around the model decides what it may do.
- Ordinary system design still applies in full.
Try it#
- Take an endpoint you maintain. If a model call of 2,000 input and 400 output tokens were added to it, what would the new p50 latency be at a TPOT of 20 ms? Which timeout breaks first?
- For a product you use, decide which of the three shapes it is and why it could not be the simpler one.
Check yourself#
- Why is a requests-per-minute limit the wrong control for a model endpoint?
- What are the two numbers that describe model latency, and what does each depend on?
- Why can a system prompt not be the place where permissions are enforced?