The idea in one minute#
Every design in this course is worked in the same six steps: requirements → numbers → architecture → failure → cost → security. The order matters. Numbers before boxes, because token arithmetic decides whether you need one API key or forty GPUs. Tools last, because a tool is an answer and you do not have the question yet.
The step that is new for AI is the numbers. You need about ten of them, and they multiply.
A picture#
flowchart LR R[":i-list-checks: <b>1 Requirements</b><br/><small>who, what, how good, how fast</small>"] --> N[":i-gauge: <b>2 Numbers</b><br/><small>tokens/s, GPUs, dollars</small>"] N --> A[":i-workflow: <b>3 Architecture</b><br/><small>boxes, arrows, state</small>"] A --> F[":i-triangle-alert: <b>4 Failure</b><br/><small>what breaks, what then</small>"] F --> C[":i-coins: <b>5 Cost</b><br/><small>per request, per month</small>"] C --> S[":i-shield-check: <b>6 Security</b><br/><small>trust boundaries, controls</small>"] S -.->|"a finding changes the design"| A class R neutral class N compute class A io class F warn class C queue class S warn
How it really works#
Step 1 — requirements#
Ask until you can fill this in. Half of bad AI designs come from skipping the quality row.
| Question | Example answer | Why it matters |
|---|---|---|
| Who uses it, and how many at once? | 20,000 employees, 2,000 at peak | Concurrency sets capacity |
| What is one unit of work? | One support ticket resolved | The thing you price and evaluate |
| Interactive or background? | Chat: interactive. Triage: background | Decides latency targets and batching |
| How good is good enough? | 90% of answers rated correct; 0 actions without approval | Sets the evaluation and the model tier |
| How fast? | First token under 1 s, full answer under 15 s | TTFT and TPOT targets |
| What may it touch? | Read the ticket system; write only drafts | The permission boundary |
| What data, and whose? | Customer messages, EU residents | Residency, retention, hosted versus self-hosted |
| What happens when it is wrong? | A human reviews before sending | Blast radius and oversight |
Step 2 — numbers#
Work in tokens. The ten numbers:
1 active users at peak U
2 requests per user per minute r
3 model calls per request k (1 for chat, 5-50 for an agent)
4 input tokens per call Tin (grows with context)
5 output tokens per call Tout
6 cached share of input c (0 to ~0.9)
7 output speed per stream 1/TPOT tokens/s
8 throughput of one GPU, for this model G output tokens/s at your latency target
9 price per million tokens, in and out Pin, Pout
10 target occupancy ~0.6 (headroom for bursts and failures)From which:
calls/s = U × r × k / 60
output tokens/s = calls/s × Tout
input tokens/s = calls/s × Tin
GPUs (self-hosted) = output tokens/s ÷ G ÷ occupancy (then add spares)
cost per request = k × (Tin × (1 − c) × Pin + Tin × c × Pcached + Tout × Pout) / 1e6Three habits make these numbers honest:
- Measure G, do not look it up. Throughput depends on the model, the engine, the quantization, the context length and your latency target. Capacity math shows how to derive it; a load test confirms it.
- Input dominates agents. An agent resends its growing history on every step, so input tokens per task grow roughly with the square of the step count. Caching and compaction are what bring that back down.
- Use the peak, not the average. Traffic for an internal tool is often five times the daily mean at 10:00 on a Monday.
Step 3 — architecture#
Draw the request path first, then state, then the control plane. For each box write one line: what it does, what it stores, what it does when the next box is slow. The building blocks are the next topic; the full shapes are in Worked Designs.
A useful rule for where state lives: the model process holds none. Conversation, task progress and memory live in stores you control, so any model replica — or a different model — can take the next call.
Step 4 — failure#
Walk every arrow and ask “what if this is slow, wrong or down?” AI adds four failure kinds that ordinary reviews miss:
| Failure | Example | Design answer |
|---|---|---|
| Slow, not down | Provider TTFT goes from 0.5 s to 12 s | Deadline per call, fallback route, shed load |
| Wrong, not failed | Valid JSON, wrong answer | Evaluations, validators, human review for high stakes |
| Runaway | An agent loops 400 times | Step and token budgets per task, enforced outside the model |
| Hijacked | A document tells the agent to email secrets | Least privilege, isolation, approval gates |
Step 5 — cost#
Compute cost per unit of work, then multiply. Compare it with what the unit is worth. Then find the two largest terms and ask what halves each: a smaller model for easy calls, caching the stable prefix, shortening output, batching background work.
Step 6 — security#
Mark every place where text you do not control enters the context, and every place where model output causes an effect. The design is safe when no path from the first set to the second passes through without a control you can name. Security by Design works this step in full.
The one-page design document#
Every design in this course fits this template. Use it for your own.
1. Problem and unit of work
2. Requirements users, quality bar, latency, data, permissions
3. Numbers the ten numbers, peak tokens/s, GPUs or spend
4. Architecture one diagram; one line per box; where state lives
5. Model choices which model for which call, and the fallback
6. Evaluation the dataset, the metric, the release gate
7. Failure modes the table above, filled in
8. Cost per unit, per month, top two terms and how to cut them
9. Security untrusted inputs, effects, controls, residual risk
10. Rollout what ships first, what is measured before the next stepCode#
The numbers of step 2 as a program. Change the constants and rerun it for any design.
// estimate.go — size an AI workload from the ten numbers.
package main
import "fmt"
type workload struct {
name string
users float64 // active at peak
reqPerUserMin float64
callsPerReq float64
tokIn, tokOut float64 // per model call
cached float64 // share of input tokens read from cache
}
func main() {
const (
gpuTokPerSec = 2500.0 // measured: output tokens/s of one GPU at the latency target
occupancy = 0.6
priceIn = 3.00 // $ per million input tokens
priceCached = 0.30
priceOut = 15.00
gpuHour = 4.00 // $ per GPU-hour
)
loads := []workload{
{"chat assistant", 2000, 0.5, 1, 3000, 350, 0.6},
{"RAG answers", 2000, 0.3, 2, 9000, 400, 0.3},
{"coding agent", 300, 0.1, 30, 40000, 600, 0.85},
}
fmt.Printf("%-16s %9s %11s %11s %6s %12s %12s\n", "workload", "calls/s", "in tok/s", "out tok/s", "GPUs", "$/request", "API $/day")
for _, w := range loads {
calls := w.users * w.reqPerUserMin * w.callsPerReq / 60
in, out := calls*w.tokIn, calls*w.tokOut
gpus := out / gpuTokPerSec / occupancy
perReq := w.callsPerReq * (w.tokIn*(1-w.cached)*priceIn + w.tokIn*w.cached*priceCached + w.tokOut*priceOut) / 1e6
reqPerDay := w.users * w.reqPerUserMin * 60 * 8 // an eight-hour peak-equivalent day
fmt.Printf("%-16s %9.1f %11.0f %11.0f %6.1f %12.4f %12.0f\n", w.name, calls, in, out, gpus, perReq, perReq*reqPerDay)
}
fmt.Printf("\nOne GPU costs $%.0f/day and serves %.0f M output tokens/day at %.0f%% occupancy.\n",
gpuHour*24, gpuTokPerSec*occupancy*86400/1e6, occupancy*100)
fmt.Println("Input tokens are counted here for price only: on your own GPUs they cost prefill time, which a load test measures.")
}Run it and notice two things: the coding agent has the fewest users and the largest bill, and its cost is almost entirely input tokens. That is the shape of every agentic workload.
Remember this#
- Six steps, always in order: requirements, numbers, architecture, failure, cost, security.
- Ten numbers size any AI workload; they multiply, so one wrong factor ruins the estimate.
- Agents are input-heavy: cost grows faster than step count.
- A design is finished when each failure and each untrusted input has a named answer.
Try it#
- Run
estimate.go. Set the coding agent’s cached share to 0. What happens to cost per request, and what does that tell you about prompt layout? - Fill in the requirements table for a product you would like to build. Which row could you not answer?
- Halve
gpuTokPerSec. Which workload’s GPU count changes most in absolute terms?
Check yourself#
- Why do numbers come before architecture?
- Why does an agent’s input-token total grow faster than its number of steps?
- Name the four failure kinds that AI adds to an ordinary failure review.