Pidoku

The Design Method

Foundations 45 min Difficulty 2/5 Lesson 02 of 03

Prerequisites What a Model Changes

The idea in one minute#

Every design in this course is worked in the same six steps: requirements → numbers → architecture → failure → cost → security. The order matters. Numbers before boxes, because token arithmetic decides whether you need one API key or forty GPUs. Tools last, because a tool is an answer and you do not have the question yet.

The step that is new for AI is the numbers. You need about ten of them, and they multiply.

A picture#

flowchart LR
  R[":i-list-checks: <b>1 Requirements</b><br/><small>who, what, how good, how fast</small>"] --> N[":i-gauge: <b>2 Numbers</b><br/><small>tokens/s, GPUs, dollars</small>"]
  N --> A[":i-workflow: <b>3 Architecture</b><br/><small>boxes, arrows, state</small>"]
  A --> F[":i-triangle-alert: <b>4 Failure</b><br/><small>what breaks, what then</small>"]
  F --> C[":i-coins: <b>5 Cost</b><br/><small>per request, per month</small>"]
  C --> S[":i-shield-check: <b>6 Security</b><br/><small>trust boundaries, controls</small>"]
  S -.->|"a finding changes the design"| A
  class R neutral
  class N compute
  class A io
  class F warn
  class C queue
  class S warn

How it really works#

Step 1 — requirements#

Ask until you can fill this in. Half of bad AI designs come from skipping the quality row.

QuestionExample answerWhy it matters
Who uses it, and how many at once?20,000 employees, 2,000 at peakConcurrency sets capacity
What is one unit of work?One support ticket resolvedThe thing you price and evaluate
Interactive or background?Chat: interactive. Triage: backgroundDecides latency targets and batching
How good is good enough?90% of answers rated correct; 0 actions without approvalSets the evaluation and the model tier
How fast?First token under 1 s, full answer under 15 sTTFT and TPOT targets
What may it touch?Read the ticket system; write only draftsThe permission boundary
What data, and whose?Customer messages, EU residentsResidency, retention, hosted versus self-hosted
What happens when it is wrong?A human reviews before sendingBlast radius and oversight

Step 2 — numbers#

Work in tokens. The ten numbers:

1  active users at peak                    U
2  requests per user per minute            r
3  model calls per request                 k     (1 for chat, 5-50 for an agent)
4  input tokens per call                   Tin   (grows with context)
5  output tokens per call                  Tout
6  cached share of input                   c     (0 to ~0.9)
7  output speed per stream                 1/TPOT tokens/s
8  throughput of one GPU, for this model   G     output tokens/s at your latency target
9  price per million tokens, in and out    Pin, Pout
10 target occupancy                        ~0.6  (headroom for bursts and failures)

From which:

calls/s            = U × r × k / 60
output tokens/s    = calls/s × Tout
input tokens/s     = calls/s × Tin
GPUs (self-hosted) = output tokens/s ÷ G ÷ occupancy       (then add spares)
cost per request   = k × (Tin × (1 − c) × Pin + Tin × c × Pcached + Tout × Pout) / 1e6

Three habits make these numbers honest:

  • Measure G, do not look it up. Throughput depends on the model, the engine, the quantization, the context length and your latency target. Capacity math shows how to derive it; a load test confirms it.
  • Input dominates agents. An agent resends its growing history on every step, so input tokens per task grow roughly with the square of the step count. Caching and compaction are what bring that back down.
  • Use the peak, not the average. Traffic for an internal tool is often five times the daily mean at 10:00 on a Monday.

Step 3 — architecture#

Draw the request path first, then state, then the control plane. For each box write one line: what it does, what it stores, what it does when the next box is slow. The building blocks are the next topic; the full shapes are in Worked Designs.

A useful rule for where state lives: the model process holds none. Conversation, task progress and memory live in stores you control, so any model replica — or a different model — can take the next call.

Step 4 — failure#

Walk every arrow and ask “what if this is slow, wrong or down?” AI adds four failure kinds that ordinary reviews miss:

FailureExampleDesign answer
Slow, not downProvider TTFT goes from 0.5 s to 12 sDeadline per call, fallback route, shed load
Wrong, not failedValid JSON, wrong answerEvaluations, validators, human review for high stakes
RunawayAn agent loops 400 timesStep and token budgets per task, enforced outside the model
HijackedA document tells the agent to email secretsLeast privilege, isolation, approval gates

Step 5 — cost#

Compute cost per unit of work, then multiply. Compare it with what the unit is worth. Then find the two largest terms and ask what halves each: a smaller model for easy calls, caching the stable prefix, shortening output, batching background work.

Step 6 — security#

Mark every place where text you do not control enters the context, and every place where model output causes an effect. The design is safe when no path from the first set to the second passes through without a control you can name. Security by Design works this step in full.

The one-page design document#

Every design in this course fits this template. Use it for your own.

1. Problem and unit of work
2. Requirements      users, quality bar, latency, data, permissions
3. Numbers           the ten numbers, peak tokens/s, GPUs or spend
4. Architecture      one diagram; one line per box; where state lives
5. Model choices     which model for which call, and the fallback
6. Evaluation        the dataset, the metric, the release gate
7. Failure modes     the table above, filled in
8. Cost              per unit, per month, top two terms and how to cut them
9. Security          untrusted inputs, effects, controls, residual risk
10. Rollout          what ships first, what is measured before the next step

Code#

The numbers of step 2 as a program. Change the constants and rerun it for any design.

Go
// estimate.go — size an AI workload from the ten numbers.
package main

import "fmt"

type workload struct {
	name          string
	users         float64 // active at peak
	reqPerUserMin float64
	callsPerReq   float64
	tokIn, tokOut float64 // per model call
	cached        float64 // share of input tokens read from cache
}

func main() {
	const (
		gpuTokPerSec = 2500.0 // measured: output tokens/s of one GPU at the latency target
		occupancy    = 0.6
		priceIn      = 3.00 // $ per million input tokens
		priceCached  = 0.30
		priceOut     = 15.00
		gpuHour      = 4.00 // $ per GPU-hour
	)
	loads := []workload{
		{"chat assistant", 2000, 0.5, 1, 3000, 350, 0.6},
		{"RAG answers", 2000, 0.3, 2, 9000, 400, 0.3},
		{"coding agent", 300, 0.1, 30, 40000, 600, 0.85},
	}
	fmt.Printf("%-16s %9s %11s %11s %6s %12s %12s\n", "workload", "calls/s", "in tok/s", "out tok/s", "GPUs", "$/request", "API $/day")
	for _, w := range loads {
		calls := w.users * w.reqPerUserMin * w.callsPerReq / 60
		in, out := calls*w.tokIn, calls*w.tokOut
		gpus := out / gpuTokPerSec / occupancy
		perReq := w.callsPerReq * (w.tokIn*(1-w.cached)*priceIn + w.tokIn*w.cached*priceCached + w.tokOut*priceOut) / 1e6
		reqPerDay := w.users * w.reqPerUserMin * 60 * 8 // an eight-hour peak-equivalent day
		fmt.Printf("%-16s %9.1f %11.0f %11.0f %6.1f %12.4f %12.0f\n", w.name, calls, in, out, gpus, perReq, perReq*reqPerDay)
	}
	fmt.Printf("\nOne GPU costs $%.0f/day and serves %.0f M output tokens/day at %.0f%% occupancy.\n",
		gpuHour*24, gpuTokPerSec*occupancy*86400/1e6, occupancy*100)
	fmt.Println("Input tokens are counted here for price only: on your own GPUs they cost prefill time, which a load test measures.")
}

Run it and notice two things: the coding agent has the fewest users and the largest bill, and its cost is almost entirely input tokens. That is the shape of every agentic workload.

Remember this#

  • Six steps, always in order: requirements, numbers, architecture, failure, cost, security.
  • Ten numbers size any AI workload; they multiply, so one wrong factor ruins the estimate.
  • Agents are input-heavy: cost grows faster than step count.
  • A design is finished when each failure and each untrusted input has a named answer.

Try it#

  1. Run estimate.go. Set the coding agent’s cached share to 0. What happens to cost per request, and what does that tell you about prompt layout?
  2. Fill in the requirements table for a product you would like to build. Which row could you not answer?
  3. Halve gpuTokPerSec. Which workload’s GPU count changes most in absolute terms?

Check yourself#

  1. Why do numbers come before architecture?
  2. Why does an agent’s input-token total grow faster than its number of steps?
  3. Name the four failure kinds that AI adds to an ordinary failure review.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom