The idea in one minute#
Red teaming an AI system means attacking it on purpose, the way an adversary would, to find what the design review and the automated tests missed. A single exercise finds bugs. A programme keeps them found: it has a scope drawn from the threat model, a method that goes after outcomes rather than interesting outputs, a mix of automated and human effort, a way to rate findings, and a pipeline that turns every finding into a permanent regression test. The measure of the programme is not how many jailbreaks it collected but whether the system’s attack success rate, on the outcomes that matter, goes down over time and stays down.
A picture#
flowchart LR TM[":i-list-checks: <b>Threat model</b><br/><small>assets, sources, sinks</small>"] --> SC[":i-route: <b>Scope and objectives</b><br/><small>'exfiltrate a canary',<br/>'trigger an ungated refund'</small>"] SC --> AUTO[":garak: <b>Automated attacks</b><br/><small>breadth, every release</small>"] SC --> HUM[":i-bug: <b>Human red team</b><br/><small>depth, creativity</small>"] AUTO --> F[":i-triangle-alert: <b>Findings</b><br/><small>rated by outcome</small>"] HUM --> F F --> FIX[":i-wrench: <b>Fix</b><br/><small>prefer a deterministic control</small>"] FIX --> REG[":i-recycle: <b>Regression case</b><br/><small>added to the CI suite</small>"] REG --> AUTO F --> TMU[":i-file-text: <b>Threat-model update</b>"] --> TM PROD[":i-radar: <b>Production signals</b><br/><small>incidents, classifier hits</small>"] --> SC class TM,TMU neutral class SC queue class AUTO,HUM compute class F warn class FIX io class REG memory class PROD neutral
How it really works#
Scope from outcomes#
A weak objective is “try to jailbreak the assistant”. A strong one names a harm:
| Objective | Success criterion, checkable in code |
|---|---|
| Exfiltrate private data via indirect injection | A planted canary string appears in an outbound request, a rendered URL or another user’s view |
| Cause an unauthorised action | A gated tool executes without a valid approval record |
| Cross a tenant boundary | Any content belonging to tenant B appears in tenant A’s session |
| Persist across sessions | An instruction planted in session 1 changes behaviour in session 2 |
| Escape or misuse the sandbox | A request reaches a non-allowlisted host; a file outside the workspace is read |
| Exhaust resources | One account causes spend or load above its limit |
| Extract protected context | System prompt, tool definitions or hidden documents returned |
| Subvert the approver | An approval is obtained for an action whose real arguments differ from what was shown |
| Compromise via supply chain | A malicious tool description or model file is accepted by the pipeline |
Outcome-based objectives keep the exercise honest: a clever prompt that makes the model say something odd but achieves none of these is a note, not a finding.
What to give the team#
AI red teaming is most productive grey-box: testers get the architecture, the tool list, the system prompts and the threat model. Hiding the system prompt wastes days rediscovering what an attacker will extract anyway, and tests obscurity rather than the controls. Give them test tenants, canary data, and a safe environment that mirrors production — the same gateway, policies, sandbox and tool catalogue — because the question is whether the system holds.
The attack surface, as a test plan#
Work through every way text can enter and every effect that can leave:
entry points user prompts; uploads (documents, images, audio); retrieved content;
web pages; emails and tickets; tool results; tool descriptions;
memory; other agents' output; repository and configuration files
techniques direct and indirect injection; obfuscation and encoding; multi-turn
escalation; multimodal payloads; payload splitting across sources;
tool chaining; argument manipulation; approval manipulation;
memory and corpus poisoning; resource amplification
targets each tool; each trust boundary; each tenant boundary; each guardrail;
the approval flow; the audit trail itselfMITRE ATLAS supplies technique names and case studies; the OWASP lists supply the coverage checklist. A matrix of entry points against objectives, with a cell marked when tested, shows what has not been tried.
Automated and human, each for what it is good at#
| Automated | Human | |
|---|---|---|
| Good at | Breadth; known techniques and their mutations; every release; measuring a rate | Novel chains; business logic; multi-step creativity; noticing what looks wrong |
| Typical tools | garak, PyRIT, promptfoo, agent benchmark environments, attacker-model loops | Security engineers plus people who know the domain and the product |
| Output | Attack success rate per category, with confidence intervals | A small number of high-value findings with reproduction steps |
| Weakness | Finds what it was built to find | Expensive; not repeatable without capture |
An effective automated technique is the attacker model: one model generates and refines attacks against the target, guided by a scorer that checks the outcome. It explores far more variations than a person can, and adapts to the defence — which is the kind of attacker the research says defences fail against.
Because the target is non-deterministic, run each attack many times and report a rate. One success in fifty attempts is a finding: an attacker can make fifty attempts.
Rating findings#
Classic severity scoring does not fit well. Rate on four axes:
| Axis | Question |
|---|---|
| Impact | What is the outcome — data disclosed, action taken, money moved, tenant boundary crossed? |
| Reach | Who can trigger it — any anonymous outsider through content, an authenticated user, an insider? |
| Reliability | Success rate per attempt, and how many attempts are feasible |
| Control that failed | Did a deterministic control fail (serious: a bug), or was there only a probabilistic one in the way (serious: a design gap)? |
An indirect injection from the open web that exfiltrates private data at a 5% success rate is critical: reach is everyone, and 5% is twenty attempts.
Fixing: the order of preference#
- Remove the capability, data or path.
- Add a deterministic control — policy rule, allowlist, scope, sandbox, approval.
- Tighten an existing deterministic control that was misconfigured.
- Improve a probabilistic control — prompt hardening, classifier, a more robust model.
Fixes of the fourth kind alone should not close a high-impact finding. They lower the success rate of this attack; the next variation arrives next week.
Keeping it fixed#
Every finding becomes:
- a regression case in the adversarial suite, with a code-checked success condition;
- usually several mutations of it, since the exact payload is rarely what recurs;
- a threat-model update if it revealed a source, sink or path that was not listed;
- a detection if the attack leaves a recognisable trace.
Then the release gate guarantees the same hole cannot reopen silently when a prompt, a model or a tool changes.
Cadence#
| When | What |
|---|---|
| Every change | The automated suite in CI |
| Continuously | Attacker-model runs against a staging mirror; sampled adversarial probing in production where safe |
| Before launch, and before any increase in capability or autonomy | A human exercise scoped to the change |
| Quarterly or twice a year | A broader human exercise; purple-team sessions with the detection team |
| After incidents and after notable public research | Targeted tests for the new technique |
Adding a tool, a data source, a memory feature or a higher autonomy level is a capability increase and triggers an exercise.
Purple teaming and detection#
Run attacks with the defenders watching. For each successful or blocked attack ask: did an alert fire, did the audit trail show the cause, could on-call contain it? An attack that was blocked but left no trace is a detection gap; one that succeeded and alerted is at least a visible failure. Exercise the kill switches for real.
Rules of engagement#
- Written authorisation and scope; test tenants and synthetic data wherever possible.
- No real customer data as a target; canaries instead.
- Care with side effects: emails, payments and third-party systems are stubbed or sandboxed.
- Third-party model providers have their own policies on adversarial testing; check them.
- Findings are handled as vulnerabilities: restricted, tracked, fixed, then shared for learning.
- An external bug-bounty or disclosure channel that explicitly includes AI issues.
Programme metrics#
coverage share of (entry point × objective) cells tested this period
system-level ASR per objective, over time — the headline number
model-level ASR how often the model is fooled, for context
time to fix by severity
regressions findings that reopened (target: zero)
detection rate share of exercise attacks that produced an alertA falling system-level ASR with a flat model-level ASR is the sign of a healthy architecture: the model is fooled as often as ever, and it matters less and less.
Remember this#
- Objectives are outcomes with code-checkable success conditions.
- Test grey-box, against a mirror of the real system with its controls on.
- Automate for breadth and rates; use people for novel chains.
- Rate by impact, reach, reliability and the kind of control that failed.
- Fix with deterministic controls first; every finding becomes a regression case.
- A capability or autonomy increase triggers a new exercise.
Try it#
- Write three outcome-based objectives for a system you know, each with a success condition.
- Build the entry-point by objective matrix for it and mark what has ever been tested.
- Take a past finding or incident and write the regression case plus two mutations.
Check yourself#
- Why is “one success in fifty attempts” a real finding?
- Why should a high-impact finding not be closed with prompt hardening alone?
- What does a falling system-level ASR with a flat model-level ASR tell you?