The idea in one minute#
A design review asks the same questions of every AI design, in the same order, and treats an unanswered question as a finding. This lesson gives you the question list — thirty questions in six groups that mirror the design method — the most common findings, and six practice problems to work alone. Use the list on your own design before anyone else sees it; it is quicker than being asked in the meeting.
A picture#
flowchart LR
D[":i-file-text: <b>Design document</b>"] --> Q1[":i-list-checks: <b>Requirements</b><br/><small>unit of work, quality bar</small>"]
Q1 --> Q2[":i-gauge: <b>Numbers</b><br/><small>tokens, peak, occupancy</small>"]
Q2 --> Q3[":i-workflow: <b>Architecture</b><br/><small>state, routes, context</small>"]
Q3 --> Q4[":i-triangle-alert: <b>Failure</b><br/><small>slow, wrong, runaway</small>"]
Q4 --> Q5[":i-coins: <b>Cost</b><br/><small>per unit, levers</small>"]
Q5 --> Q6[":i-shield-check: <b>Security</b><br/><small>sources, sinks, blast radius</small>"]
Q6 --> OUT{":i-scale: <b>Outcome</b>"}
OUT -->|"no open blockers"| GO[":i-check: <b>Build the first stage</b>"]
OUT -->|"blockers"| BACK[":i-recycle: <b>Revise</b>"] --> D
class D neutral
class Q1,Q2,Q3 io
class Q4,Q6 warn
class Q5,OUT queue
class GO memory
class BACK neutralHow it really works#
The thirty questions#
Requirements
- What is one unit of work, and what is it worth?
- What is the quality bar, as a number on a named dataset?
- Which calls are interactive and which can wait?
- What may the system read, and what may it change?
- What happens, concretely, when it is wrong?
Numbers
- What are peak input and output tokens per second, per model?
- How many model calls per unit of work, and how does context grow across them?
- What cache hit rate is assumed, and what makes it true?
- For self-hosted models: what measured goodput per replica, at what latency target?
- What occupancy is planned, and what is the capacity on the bad day?
Architecture
- Is there one gateway in front of every model, and does any code bypass it?
- Where does each kind of state live, and can any model replica serve the next call?
- What exactly is in the context for each call, in what order, with what budget?
- Which model serves which call site, and what evidence chose it?
- Is every behaviour-changing artifact versioned and gated by evaluation?
Failure
- What are the first-token and total deadlines, and what happens when each passes?
- What is the fallback for each model route, and has its prompt been evaluated?
- What is shed first under overload?
- What stops a runaway loop, and where is that enforced?
- Does a crash or deploy lose task progress or repeat a side effect?
Cost
- What does one unit of work cost, and what are the two largest terms?
- What halves each of those terms?
- Who sees the bill, per tenant, and what enforces a budget?
Security
- Which text entering any context is controlled by someone else?
- What can model output cause, classified as read, write, send, execute?
- Does any single context hold private data, untrusted content and an outbound channel?
- Whose authority does each tool call carry, how is it scoped, how long does it live?
- Where does model-written code run, and what can it reach?
- What is the worst outcome of a fully manipulated agent, in one sentence?
- How would you know it had happened?
The findings that recur#
| Finding | Seen as | Fix |
|---|---|---|
| No evaluation | “We tested it and it looked good” | A dataset and a gate before anything else |
| Requests counted, not tokens | Capacity plan in requests per second | Redo the numbers in tokens, split input and output |
| State in the process | Conversation or task progress in memory | Move it to a store; make workers disposable |
| Cache-hostile prompts | Timestamp or user name at the top of the prompt | Stable prefix first; measure hit rate |
| Averages, not peaks | Sized for the daily mean | Size for peak, at target occupancy, with a replica down |
| Fallback never exercised | A second provider configured, never evaluated | Put the fallback path in the evaluation suite and test it monthly |
| Budgets in the prompt | “Do not take more than 20 steps” | Enforce in the harness |
| Autoscaling on GPU utilisation | HPA on a utilisation metric | Scale on queue wait and KV pressure |
| Permissions after retrieval | Results trimmed post-search | Filter inside the search; fail closed |
| Agent with a service account | One powerful key for all users | Delegated, scoped, short-lived tokens |
| The trifecta, unnoticed | A helpful new tool added to an agent that reads untrusted text | Redo the source–sink analysis on every tool addition |
| Open egress from sandboxes | “It needs the internet for packages” | Allowlist through a proxy |
| Approval for everything | Users click through | Gate only irreversible and external actions |
| No rollback path for prompts | Prompt edited in a console | Version control, canary, one-step revert |
Scoring a design#
A review ends with one of three outcomes:
- Blockers — must change before building: no evaluation plan, an unmitigated trifecta, an unbounded loop, lost progress on crash, permissions enforced by instruction.
- Risks — accepted knowingly, with an owner: a single provider, an untested scale point, a manual step.
- Notes — improvements for later.
Record them in the design document. A design with written, owned risks is in better shape than one with no findings because nobody looked.
Practice problems#
Work each through the six steps and the one-page template from The Design Method. Each takes about an hour. The line after each names the part most people get wrong.
- A code-review agent for a 2,000-engineer company: comments on every pull request within five minutes. Untrusted input: the diff itself, including contributions from outside.
- A customer-support agent that can look up orders and issue refunds up to a limit. Refund is an irreversible effect triggered by customer-written text.
- A meeting assistant that transcribes calls, writes summaries and files follow-up tasks. Anything said in the meeting — or shown on a shared screen — reaches the context.
- A self-hosted chat service for a hospital, on premises, 3,000 clinicians. Capacity on a fixed GPU budget; no overflow route; data residency.
- A research agent that browses the web and writes reports with citations. What it may hold in memory, and what its browser sessions are logged into.
- A multi-tenant agent-hosting product: customers upload agents that run on your infrastructure. Hostile tenants; isolation of execution, caches and model access.
For problem 6, the security treatment is worked in Securing an Agent Platform.
Presenting a design#
Whether in a review meeting or an interview, the order that works:
- Restate the problem and the unit of work in two sentences.
- Ask the requirement questions you cannot assume — quality bar, latency, permissions.
- Do the numbers aloud, in tokens. State assumptions; round freely.
- Draw the request path, then where state lives, then the control plane.
- Walk one request through the diagram.
- Break it: slow, wrong, runaway, hijacked. Say what happens for each.
- Price it per unit of work and name the biggest lever.
- Mark the trust boundaries and state the blast radius.
- Say what you would build first.
Naming tools is the least important part. A reviewer wants to know that you can state what a box must do and how it fails; which project fills it follows from that.
Remember this#
- Thirty questions in six groups; an unanswered question is a finding.
- The recurring blockers: no evaluation, state in the process, unbounded loops, permissions by instruction, an unnoticed trifecta.
- Outcomes are blockers, owned risks and notes — written down.
- Present in the order of the method; tools last.
Try it#
- Run the thirty questions against a design you have shipped. Count the ones you cannot answer today.
- Work practice problem 2 on one page.
- Exchange designs with a colleague and review each other’s with the findings table.
Check yourself#
- Which five findings should always be blockers?
- Why is an untested fallback route a finding even though it “exists”?
- In what order would you present a design, and why are the numbers before the diagram?
You have finished the course. Continue with AI Security Engineering for the security half in depth, or go down a layer into Inference Engineering to see what happens inside the serving plane.