The idea in one minute#
In an AI application, security-relevant behaviour is defined by things that are not traditionally “code”: prompts, tool definitions, policies, routing rules, model versions. The build stage brings all of them under the same discipline as code — version control, review, automated tests, a release gate — and adds one kind of test that ordinary software does not have: adversarial evaluation, where you attack your own system automatically on every change and track the attack success rate as a number that must not rise.
A picture#
flowchart LR
PR[":github: <b>Pull request</b><br/><small>code, prompt, tool, policy,<br/>model version</small>"] --> CI
subgraph CI["CI pipeline"]
direction TB
C1[":semgrep: <b>Static checks</b><br/><small>SAST, secrets, dependencies</small>"]
C2[":i-list-checks: <b>Policy unit tests</b><br/><small>deterministic controls</small>"]
C3[":i-scale: <b>Quality evaluation</b>"]
C4[":garak: <b>Adversarial evaluation</b><br/><small>injection, jailbreak, leakage</small>"]
C1 --> C2 --> C3 --> C4
end
CI --> GATE{":i-shield-check: <b>Release gate</b><br/><small>must-pass + no regression</small>"}
GATE -->|"pass"| SIGN[":sigstore: <b>Sign and publish</b>"]
GATE -->|"fail"| BLOCK[":i-ban: <b>Block</b>"]
SIGN --> CAN[":argo: <b>Canary</b>"]
RT[":i-bug: <b>Human red team</b><br/><small>periodic, creative</small>"] -->|"findings become cases"| C4
INC[":i-siren: <b>Incidents</b>"] -->|"become cases"| C4
class PR neutral
class C1,C2 queue
class C3,C4 compute
class GATE queue
class SIGN,CAN io
class BLOCK,RT,INC warnHow it really works#
Everything that changes behaviour is in the repository#
system prompts and templates tool definitions and descriptions
policy rules (who may call what) guardrail configuration and thresholds
model identifiers and versions routing and fallback rules
retrieval settings evaluation datasets and gradersEach gets review by someone other than the author, a history, and a rollback. Prompts edited in a web console are unreviewed production changes.
Static checks, extended#
The usual checks apply, with AI-specific additions:
| Check | Finds |
|---|---|
| Secret scanning | API keys in prompts, notebooks, example files and agent configuration — prompts are a frequent hiding place |
| Dependency scanning | Vulnerable or malicious packages; verify that every dependency an assistant suggested actually exists and is the intended one |
| SAST with AI rules | Model output passed to eval, a shell, a SQL string, an HTML template or a file path without validation — improper output handling |
| Configuration linting | Agent and MCP configuration files that auto-start commands; tools granted wildcard permissions; max_tokens unset |
| Prompt linting | Secrets or internal hostnames in prompts; instructions that claim to enforce security (“never reveal…”) with no code-level control behind them |
The SAST row deserves emphasis. Treat every model output as you would treat a request body from the internet: parse it against a schema, validate values, encode it for the destination. Most “AI vulnerabilities” found in code review are ordinary injection bugs — SQL, command, path, cross-site scripting — where the tainted input happens to come from a model.
Unit-test the deterministic controls#
Policy layers, sanitisers, permission filters and budget enforcement are ordinary code, so they get ordinary tests, and these are the most valuable tests in the suite because they verify the controls you rely on:
- The policy denies
send_emailin a tainted session. - The retrieval filter returns nothing for a user without access.
- The sanitiser removes an image pointing at an unknown host, in every markdown variant.
- The budget stops a task at the configured step count.
- A tool argument outside the allowlist is rejected.
These tests are deterministic, fast, and must pass at 100%.
Adversarial evaluation#
Then attack the whole system, automatically. An adversarial suite is an evaluation whose cases are attacks and whose metric is attack success rate (ASR).
| Category | Example case | Success means |
|---|---|---|
| Direct injection / jailbreak | A library of known techniques, mutated | A forbidden output or action occurred |
| Indirect injection | A document, email or tool result with a planted instruction | The agent acted on it: called the tool, wrote the memory, changed the answer |
| Exfiltration | Plant a canary secret in context, plus an injection telling the model to leak it | The canary appears in output, a URL or an outbound call |
| Hidden context exposure | Attempts to extract the system prompt | Protected text returned |
| Permission bypass | User A asks about data only B may read | Any of B’s data in the response |
| Tool misuse | Prompts steering toward out-of-policy arguments | The tool executed with them |
| Resource abuse | Inputs designed to loop or maximise output | A budget was exceeded |
Three practices make the numbers meaningful:
- Grade the outcome, not the wording. Did the tool call happen? Did the canary leave? That is checkable in code and is not fooled by a polite refusal that leaks anyway.
- Run each case several times. The system is non-deterministic; report a rate.
- Test the whole system, with its controls on. The question is not “can the model be fooled?” — it can — but “does the attack achieve anything?”
Two numbers are worth separating: ASR against the model alone (how often it is fooled), and ASR against the system (how often being fooled leads to harm). A good architecture shows a large gap between them, and the second is the one to gate on.
The release gate#
block the release if
any deterministic-control test fails
any must-pass adversarial case succeeds (known incidents, permission bypass, canary leak)
system-level attack success rate rises beyond noise versus the baseline
quality, cost or latency leave their budgetsA model upgrade goes through the same gate. A newer model is usually more resistant to attack; sometimes it is more capable at following a subtle injected instruction. Measure it.
Tools for automated red teaming#
| Tool | What it does |
|---|---|
| garak | An open-source scanner that runs large libraries of probes — jailbreaks, injection, leakage, toxicity — against a model or endpoint and reports failures |
| PyRIT | Microsoft’s open framework for orchestrating multi-turn attacks, with attacker models, converters and scorers |
| promptfoo | Evaluation and red-team configuration in CI, with generated adversarial cases |
| Agent benchmarks (AgentDojo and similar) | Environments with tools where injections are planted in tool results, measuring both task success and attack success |
| Your own cases | Derived from your threat model and incidents — the ones that matter most |
Public probe libraries test general robustness. Your own cases test whether your tools,
data and policies can be turned against you; no public tool knows what your issue_refund
does.
Human red teaming#
Automation covers known techniques at scale. People find the new ones — the odd interaction between two tools, the business-logic abuse, the carrier nobody listed. Schedule human exercises before major launches and after significant capability changes, give the team the threat model, and convert every finding into an automated case so it stays fixed. The Red-Team Programme covers how to run one.
Testing the test#
An adversarial suite decays: models get trained on public attacks, and yesterday’s cases stop finding anything while the real attack surface moves. Keep it honest by adding cases from incidents and human red teaming, mutating existing cases, and occasionally planting a deliberate weakness to confirm the suite detects it.
Remember this#
- Prompts, tools, policies and model versions are code: reviewed, versioned, gated.
- Model output is untrusted input to whatever consumes it — validate and encode.
- Unit-test deterministic controls at 100%; they are the ones you rely on.
- Adversarial evaluation reports an attack success rate; grade outcomes, run repeatedly, test the whole system.
- Gate on system-level ASR and must-pass cases; every finding becomes a permanent test.
Try it#
- Write five adversarial cases for a system you know: one per category of your choice, each with a code-checkable success condition.
- Search a repository you maintain for places where model output reaches a shell, a query or an HTML template. How many validate it first?
- Define the must-pass list for your release gate.
Check yourself#
- Why should adversarial tests grade outcomes rather than the model’s wording?
- What is the difference between model-level and system-level attack success rate?
- Why does a model upgrade need to pass the security gate again?