The idea in one minute#
You cannot unit-test a probabilistic component with assert equal. You evaluate it: run
it on a fixed set of cases, score each output, and track the pass rate. An evaluation is to an
AI system what a test suite is to ordinary code — the thing that lets you change a prompt, a
model or a retrieval setting and know whether you made it better or worse. Teams without one
ship on intuition, and intuition does not notice a 4% regression.
Evaluation happens in two places: offline, before a change ships, and online, on real traffic after it ships.
A picture#
flowchart LR
subgraph OFF["Offline: before release"]
direction LR
DS[(":i-database: <b>Dataset</b><br/><small>cases + expected</small>")] --> RUN[":i-bot: <b>Run the system</b>"]
RUN --> SC[":i-scale: <b>Graders</b><br/><small>code, model judge, human</small>"]
SC --> GATE{":i-list-checks: <b>Gate</b><br/><small>no worse than baseline?</small>"}
end
GATE -->|"pass"| REL[":argo: <b>Rollout</b><br/><small>canary, then all</small>"]
GATE -->|"fail"| FIX[":i-wrench: Fix and rerun"]
subgraph ON["Online: after release"]
direction LR
TR[":opentelemetry: <b>Traces</b>"] --> SIG[":i-activity: <b>Signals</b><br/><small>thumbs, retries, edits, judge samples</small>"]
SIG --> REV[":i-eye: <b>Review failures</b>"]
end
REL --> TR
REV -->|"new cases"| DS
class DS memory
class RUN,SC compute
class GATE queue
class REL,TR io
class FIX,REV warn
class SIG neutralThe arrow from “review failures” back to the dataset is the whole method. Every production failure becomes a test case, so the same failure cannot ship twice.
How it really works#
The dataset#
Start with 30 to 50 real cases, not 5,000 synthetic ones. Each case has an input, whatever context the system needs, and a way to judge the output. Cover:
- Typical requests, in the proportions you actually see.
- Edge cases: empty input, very long input, the wrong language, ambiguous questions.
- Adversarial cases: attempts to break the rules — these double as security regression tests.
- Past failures, each added on the day it was found.
Keep a held-out portion that nobody tunes prompts against, or the suite stops measuring anything.
Three kinds of grader#
| Grader | Use for | Strength | Weakness |
|---|---|---|---|
| Code | Valid JSON, required fields, exact answers, forbidden strings, tool-call correctness, latency, cost | Fast, free, deterministic | Cannot judge meaning |
| Model judge | Faithfulness to sources, helpfulness, tone, rubric scoring | Scales to open-ended output | Is itself a model: must be validated against humans, and can be fooled |
| Human | Building and calibrating the others; high-stakes samples | Ground truth | Slow, costly, inconsistent between people |
Use code wherever code can decide. Use a judge model with a specific rubric (“Does the answer contain a claim not supported by the provided sources? Answer yes or no, then quote it”) and check its agreement with human labels on a sample before trusting its numbers. A judge that agrees with people 70% of the time is a random number generator with a good vocabulary.
What to measure, by system shape#
| Shape | Measure |
|---|---|
| Single call | Accuracy against labels; schema validity; cost and latency per call |
| Retrieval-augmented | Retrieval recall@k and answer faithfulness, separately; citation correctness; “I do not know” when the answer is absent |
| Agentic | Task success judged on the end state (was the ticket closed correctly?), not on the path; steps and tokens per task; whether any forbidden action was attempted |
For agents, grading the end state matters: there are many good paths to a correct outcome, and a grader that demands one specific sequence of tool calls punishes better solutions.
Statistics you cannot skip#
Pass rates on small datasets are noisy. With 50 cases, a move from 80% to 84% is two cases and means nothing. Three habits:
- Report a confidence interval with every pass rate.
- Run each case several times; the system is non-deterministic and a case can pass one run in three.
- Compare a change to the baseline on the same cases and look at which cases flipped, not only at the totals.
The release gate#
Every change that can alter behaviour runs the evaluation in CI: prompt edits, model version changes, retrieval parameters, tool descriptions, routing rules. The gate has three parts:
- Quality is not worse than the baseline beyond noise.
- No case tagged must-pass fails — safety rules, known past incidents.
- Cost and latency stay inside budget.
Then roll out as you would any risky change: a canary on a few percent of traffic, compared to the baseline on online signals, before the rest.
Online evaluation#
Offline suites cannot contain what you have not imagined. In production:
- Implicit signals are the richest: did the user retry, rephrase, copy the answer, edit the draft, abandon the session?
- Explicit feedback is sparse and biased toward the unhappy; still worth collecting.
- Sampled judging: run the model judge on a small percentage of live traffic and chart the score per model version and prompt version.
- Review sessions: read actual traces every week. Nothing replaces reading what the system did. Quality and evals covers the telemetry side.
What fills the box in 2026#
Tracing-and-evaluation platforms (Langfuse, LangSmith, Braintrust, Arize Phoenix, MLflow, Weights & Biases Weave) store datasets, run graders and compare versions. Open benchmark harnesses exist for agents and for security testing. All of them are optional; a directory of JSON cases and the fifty-line harness below is a legitimate start, and better than none.
Code#
A minimal evaluation harness: cases, a code grader, repeated runs, and a confidence interval.
// eval.go — run cases several times, grade with code, report a pass rate with an interval.
package main
import (
"fmt"
"math"
"math/rand"
"strings"
)
type testCase struct {
name, input string
mustContain []string
mustNot []string
}
// system stands in for the real pipeline. It is deliberately flaky on one case.
func system(rng *rand.Rand, input string) string {
switch {
case strings.Contains(input, "refund"):
if rng.Float64() < 0.25 {
return "Sure, I have refunded you." // the failure: acting without the policy check
}
return "Refunds are available within 30 days. Shall I check your order?"
case strings.Contains(input, "password"):
return "I can't share credentials. Use the reset link in settings."
default:
return "Here is the information you asked for."
}
}
func grade(c testCase, out string) bool {
for _, s := range c.mustContain {
if !strings.Contains(out, s) {
return false
}
}
for _, s := range c.mustNot {
if strings.Contains(out, s) {
return false
}
}
return true
}
// wilson returns a 95% confidence interval for a pass rate.
func wilson(pass, n float64) (lo, hi float64) {
const z = 1.96
p := pass / n
centre := (p + z*z/(2*n)) / (1 + z*z/n)
half := z * math.Sqrt(p*(1-p)/n+z*z/(4*n*n)) / (1 + z*z/n)
return centre - half, centre + half
}
func main() {
cases := []testCase{
{"refund policy", "I want a refund", []string{"30 days"}, []string{"have refunded"}},
{"no secrets", "what is the admin password", []string{"can't share"}, nil},
{"general", "opening hours?", []string{"information"}, nil},
}
const runs = 20
rng := rand.New(rand.NewSource(7))
var pass, total float64
for _, c := range cases {
ok := 0
for i := 0; i < runs; i++ {
if grade(c, system(rng, c.input)) {
ok++
}
}
fmt.Printf("%-14s %2d/%d\n", c.name, ok, runs)
pass += float64(ok)
total += runs
}
lo, hi := wilson(pass, total)
fmt.Printf("\noverall %.1f%% (95%% interval %.1f%% to %.1f%%)\n", 100*pass/total, 100*lo, 100*hi)
fmt.Println("A single run per case would have reported either 100% or 67%, depending on luck.")
}Remember this#
- An evaluation is the test suite of an AI system. No evaluation, no safe change.
- Start with a few dozen real cases; add every production failure.
- Grade with code where possible; validate judge models against humans.
- Grade agents on end state, and run cases more than once.
- Gate releases on quality, must-pass cases, and cost and latency together.
Try it#
- Run
eval.go, then changerunsto 1 and the seed a few times. How often does the flaky case hide? - Write five evaluation cases for a feature you maintain, including one adversarial case.
- Write a yes/no rubric a judge model could apply to “is this answer supported by its sources?”.
Check yourself#
- Why should a judge model’s output never be trusted before it is compared with human labels?
- Why grade an agent on its end state rather than its sequence of steps?
- Which changes must pass the evaluation gate besides model upgrades?