Pidoku

Reliability Patterns

Basic 45 min Difficulty 3/5 Lesson 06 of 06

Prerequisites Model Access and Gateways

The idea in one minute#

A model dependency fails in ways an ordinary dependency does not. It gets slow without failing, it is rate-limited by tokens, it returns well-formed nonsense, and one request can cost a thousand times more than another. The classic patterns — deadlines, retries, circuit breakers, fallbacks, load shedding — all still apply, but each needs retuning for requests that last thirty seconds and stream as they go. And agents add one new requirement: a budget that stops a loop the model will not stop itself.

A picture#

flowchart LR
  REQ[":i-user: <b>Request</b>"] --> ADM{":i-funnel: <b>Admit?</b><br/><small>queue depth, budget</small>"}
  ADM -->|"over limit"| SHED[":i-ban: <b>Shed</b><br/><small>429 + retry-after</small>"]
  ADM --> DL[":i-clock: <b>Deadline</b><br/><small>first token: 3 s<br/>total: 60 s</small>"]
  DL --> CB{":i-zap: <b>Breaker</b><br/><small>route healthy?</small>"}
  CB -->|"closed"| P[":anthropic: <b>Primary model</b>"]
  CB -->|"open"| F[":vllm: <b>Fallback model</b>"]
  P -->|"timeout or 5xx"| RT[":i-recycle: <b>Retry</b><br/><small>with a budget</small>"]
  RT --> F
  P --> V[":i-shield-check: <b>Validate output</b>"]
  F --> V
  V -->|"invalid"| DEG[":i-triangle-alert: <b>Degrade</b><br/><small>cached, simpler, or honest error</small>"]
  V --> OK[":i-check: <b>Respond</b>"]
  class REQ,OK neutral
  class ADM,CB,DL queue
  class SHED,DEG,RT warn
  class P,F compute
  class V io

How it really works#

Deadlines, in two parts#

One timeout is not enough for a stream. Use two:

  • Time to first token. If nothing has arrived in a few seconds the request is probably stuck in a queue; cancel and go elsewhere while switching is still invisible to the user.
  • Total time, plus a stall timer between tokens, to catch a stream that starts and then hangs.

Propagate the remaining deadline downstream on every call, and cancel work when the client disconnects. An abandoned generation still occupies a GPU, or still bills you, until someone cancels it.

Retries, with a budget#

Retries turn a brief overload into a sustained one: every client tries again at once and the load multiplies. Rules:

  • Retry only what is retryable: timeouts before first token, 429, 5xx. Never a 400.
  • Exponential backoff with jitter; honour retry-after.
  • A retry budget: retries may add at most about 10% to total traffic. Past that, fail fast.
  • Retry at one layer. Three layers each retrying three times is 27 attempts.
  • After the first token has reached the user, do not retry; you cannot un-send text.

Model calls with side effects — tool calls that write — need idempotency keys so that a retry does not perform the action twice.

Circuit breakers#

A breaker watches a route’s recent results. After enough failures it opens and requests skip that route immediately instead of waiting for a timeout. After a pause it lets a few probes through (half-open) and closes again if they succeed. For model routes, count slow first tokens as failures: a provider at 20-second TTFT is down for interactive purposes even though every request “succeeds”.

Fallback and degradation#

Decide the degraded modes in the design, not during the incident:

LevelBehaviour
1Another provider or region serving the same model
2A smaller or older model, with a tested prompt
3Reduced function: answer from retrieved text only, shorter output, no tools
4A cached or canned response, clearly labelled
5An honest error with a retry time

A fallback model is a different model: its prompt must be in the evaluation suite too, or the fallback path ships untested. And fallback capacity has to exist — a provider outage sends everyone’s traffic to the same alternatives at once.

Admission control and backpressure#

When demand exceeds capacity, something must be refused, and it is far better to choose what. Reject early, with 429 and retry-after, based on queue depth or estimated wait — not after the request has waited 30 seconds and then timed out. Give interactive traffic priority over batch, and paying tiers over free ones. On your own GPUs, the signals to admit on are queue depth and KV-cache pressure, not GPU utilisation; see queues and admission.

Output validation#

A response that parses is not a response that is right, but one that does not parse is certainly wrong. Validate structure in code, repair once (send the error back to the model), then degrade. Count repairs: a rising repair rate is an early signal that a model or prompt change has gone wrong.

Budgets for agents#

A loop has no natural end. Enforce, in the harness and never in the prompt:

per task     max steps            e.g. 50
             max total tokens     e.g. 2,000,000
             max wall-clock time  e.g. 30 min
             max spend            e.g. $5
per tool     max calls per task; rate limit per minute
globally     concurrent tasks per tenant

Add a no-progress detector — the same tool call with the same arguments three times is a loop — and a human-visible way to stop a task. When a budget is hit, the task ends in a defined state with its progress saved, not a stack trace.

The failure review#

For each arrow in a design, write one line for each of these:

slow          what is the deadline, and what happens when it passes?
down          where does traffic go?
wrong         how is bad output detected, and what then?
overloaded    what is shed first?
runaway       which budget stops it?
duplicated    is the operation idempotent?

Code#

A circuit breaker that treats slow first tokens as failures.

Go
// breaker.go — a circuit breaker for a model route, counting slow first tokens as failures.
package main

import (
	"errors"
	"fmt"
	"time"
)

type state int

const (
	closed state = iota
	open
	halfOpen
)

func (s state) String() string { return [...]string{"closed", "open", "half-open"}[s] }

type breaker struct {
	state     state
	failures  int
	threshold int
	openedAt  time.Time
	cooldown  time.Duration
	slowTTFT  time.Duration
}

var errOpen = errors.New("circuit open: skip this route")

// do runs call unless the breaker is open, and records the outcome.
func (b *breaker) do(now time.Time, call func() (ttft time.Duration, err error)) error {
	if b.state == open {
		if now.Sub(b.openedAt) < b.cooldown {
			return errOpen
		}
		b.state = halfOpen // let one probe through
	}
	ttft, err := call()
	if err != nil || ttft > b.slowTTFT {
		b.failures++
		if b.state == halfOpen || b.failures >= b.threshold {
			b.state, b.openedAt = open, now
		}
		if err == nil {
			err = fmt.Errorf("first token took %v", ttft)
		}
		return err
	}
	b.state, b.failures = closed, 0
	return nil
}

func main() {
	b := &breaker{threshold: 3, cooldown: 10 * time.Second, slowTTFT: 3 * time.Second}
	now := time.Unix(0, 0)
	// TTFT observed on each attempt, one per second: healthy, then a brown-out, then recovery.
	ttfts := []float64{0.4, 0.5, 9, 12, 15, 14, 13, 11, 10, 9, 8, 7, 6, 5, 4, 0.5, 0.4, 0.4}
	for i, s := range ttfts {
		now = now.Add(time.Second)
		err := b.do(now, func() (time.Duration, error) { return time.Duration(s * float64(time.Second)), nil })
		route := "primary"
		if errors.Is(err, errOpen) {
			route = "FALLBACK"
		}
		fmt.Printf("t=%2ds  ttft=%4.1fs  breaker=%-9s  served by %-8s %v\n", i+1, s, b.state, route, errText(err))
	}
}

func errText(err error) string {
	if err == nil {
		return ""
	}
	return "(" + err.Error() + ")"
}

During the brown-out only three requests wait for a slow first token; the rest go straight to the fallback, and the breaker closes again by itself once the primary recovers.

Remember this#

  • Two deadlines for a stream: first token, and total with a stall timer.
  • Retry at one layer, with backoff, jitter and a budget; never after the first token.
  • Slow is a failure. Breakers should count it.
  • Degraded modes are designed and evaluated in advance.
  • Agent budgets — steps, tokens, time, money — are enforced by the harness.

Try it#

  1. Run breaker.go. Lower slowTTFT to 1 s; what happens during normal operation?
  2. Write the six-line failure review for the arrow between your application and its model provider.
  3. Choose budgets for a coding agent and for a customer-support agent. Why do they differ?

Check yourself#

  1. Why do you need both a first-token deadline and a total deadline?
  2. How do retries make an overload worse, and what limits the damage?
  3. Why must an agent’s step budget be enforced outside the model?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom