Pidoku

Build and Test

Intermediate 45 min Difficulty 3/5 Lesson 04 of 07

Prerequisites The Lifecycle Map, How Attacks Work

The idea in one minute#

In an AI application, security-relevant behaviour is defined by things that are not traditionally “code”: prompts, tool definitions, policies, routing rules, model versions. The build stage brings all of them under the same discipline as code — version control, review, automated tests, a release gate — and adds one kind of test that ordinary software does not have: adversarial evaluation, where you attack your own system automatically on every change and track the attack success rate as a number that must not rise.

A picture#

flowchart LR
  PR[":github: <b>Pull request</b><br/><small>code, prompt, tool, policy,<br/>model version</small>"] --> CI
  subgraph CI["CI pipeline"]
    direction TB
    C1[":semgrep: <b>Static checks</b><br/><small>SAST, secrets, dependencies</small>"]
    C2[":i-list-checks: <b>Policy unit tests</b><br/><small>deterministic controls</small>"]
    C3[":i-scale: <b>Quality evaluation</b>"]
    C4[":garak: <b>Adversarial evaluation</b><br/><small>injection, jailbreak, leakage</small>"]
    C1 --> C2 --> C3 --> C4
  end
  CI --> GATE{":i-shield-check: <b>Release gate</b><br/><small>must-pass + no regression</small>"}
  GATE -->|"pass"| SIGN[":sigstore: <b>Sign and publish</b>"]
  GATE -->|"fail"| BLOCK[":i-ban: <b>Block</b>"]
  SIGN --> CAN[":argo: <b>Canary</b>"]
  RT[":i-bug: <b>Human red team</b><br/><small>periodic, creative</small>"] -->|"findings become cases"| C4
  INC[":i-siren: <b>Incidents</b>"] -->|"become cases"| C4
  class PR neutral
  class C1,C2 queue
  class C3,C4 compute
  class GATE queue
  class SIGN,CAN io
  class BLOCK,RT,INC warn

How it really works#

Everything that changes behaviour is in the repository#

system prompts and templates        tool definitions and descriptions
policy rules (who may call what)    guardrail configuration and thresholds
model identifiers and versions      routing and fallback rules
retrieval settings                  evaluation datasets and graders

Each gets review by someone other than the author, a history, and a rollback. Prompts edited in a web console are unreviewed production changes.

Static checks, extended#

The usual checks apply, with AI-specific additions:

CheckFinds
Secret scanningAPI keys in prompts, notebooks, example files and agent configuration — prompts are a frequent hiding place
Dependency scanningVulnerable or malicious packages; verify that every dependency an assistant suggested actually exists and is the intended one
SAST with AI rulesModel output passed to eval, a shell, a SQL string, an HTML template or a file path without validation — improper output handling
Configuration lintingAgent and MCP configuration files that auto-start commands; tools granted wildcard permissions; max_tokens unset
Prompt lintingSecrets or internal hostnames in prompts; instructions that claim to enforce security (“never reveal…”) with no code-level control behind them

The SAST row deserves emphasis. Treat every model output as you would treat a request body from the internet: parse it against a schema, validate values, encode it for the destination. Most “AI vulnerabilities” found in code review are ordinary injection bugs — SQL, command, path, cross-site scripting — where the tainted input happens to come from a model.

Unit-test the deterministic controls#

Policy layers, sanitisers, permission filters and budget enforcement are ordinary code, so they get ordinary tests, and these are the most valuable tests in the suite because they verify the controls you rely on:

  • The policy denies send_email in a tainted session.
  • The retrieval filter returns nothing for a user without access.
  • The sanitiser removes an image pointing at an unknown host, in every markdown variant.
  • The budget stops a task at the configured step count.
  • A tool argument outside the allowlist is rejected.

These tests are deterministic, fast, and must pass at 100%.

Adversarial evaluation#

Then attack the whole system, automatically. An adversarial suite is an evaluation whose cases are attacks and whose metric is attack success rate (ASR).

CategoryExample caseSuccess means
Direct injection / jailbreakA library of known techniques, mutatedA forbidden output or action occurred
Indirect injectionA document, email or tool result with a planted instructionThe agent acted on it: called the tool, wrote the memory, changed the answer
ExfiltrationPlant a canary secret in context, plus an injection telling the model to leak itThe canary appears in output, a URL or an outbound call
Hidden context exposureAttempts to extract the system promptProtected text returned
Permission bypassUser A asks about data only B may readAny of B’s data in the response
Tool misusePrompts steering toward out-of-policy argumentsThe tool executed with them
Resource abuseInputs designed to loop or maximise outputA budget was exceeded

Three practices make the numbers meaningful:

  • Grade the outcome, not the wording. Did the tool call happen? Did the canary leave? That is checkable in code and is not fooled by a polite refusal that leaks anyway.
  • Run each case several times. The system is non-deterministic; report a rate.
  • Test the whole system, with its controls on. The question is not “can the model be fooled?” — it can — but “does the attack achieve anything?”

Two numbers are worth separating: ASR against the model alone (how often it is fooled), and ASR against the system (how often being fooled leads to harm). A good architecture shows a large gap between them, and the second is the one to gate on.

The release gate#

block the release if
  any deterministic-control test fails
  any must-pass adversarial case succeeds     (known incidents, permission bypass, canary leak)
  system-level attack success rate rises beyond noise versus the baseline
  quality, cost or latency leave their budgets

A model upgrade goes through the same gate. A newer model is usually more resistant to attack; sometimes it is more capable at following a subtle injected instruction. Measure it.

Tools for automated red teaming#

ToolWhat it does
garakAn open-source scanner that runs large libraries of probes — jailbreaks, injection, leakage, toxicity — against a model or endpoint and reports failures
PyRITMicrosoft’s open framework for orchestrating multi-turn attacks, with attacker models, converters and scorers
promptfooEvaluation and red-team configuration in CI, with generated adversarial cases
Agent benchmarks (AgentDojo and similar)Environments with tools where injections are planted in tool results, measuring both task success and attack success
Your own casesDerived from your threat model and incidents — the ones that matter most

Public probe libraries test general robustness. Your own cases test whether your tools, data and policies can be turned against you; no public tool knows what your issue_refund does.

Human red teaming#

Automation covers known techniques at scale. People find the new ones — the odd interaction between two tools, the business-logic abuse, the carrier nobody listed. Schedule human exercises before major launches and after significant capability changes, give the team the threat model, and convert every finding into an automated case so it stays fixed. The Red-Team Programme covers how to run one.

Testing the test#

An adversarial suite decays: models get trained on public attacks, and yesterday’s cases stop finding anything while the real attack surface moves. Keep it honest by adding cases from incidents and human red teaming, mutating existing cases, and occasionally planting a deliberate weakness to confirm the suite detects it.

Remember this#

  • Prompts, tools, policies and model versions are code: reviewed, versioned, gated.
  • Model output is untrusted input to whatever consumes it — validate and encode.
  • Unit-test deterministic controls at 100%; they are the ones you rely on.
  • Adversarial evaluation reports an attack success rate; grade outcomes, run repeatedly, test the whole system.
  • Gate on system-level ASR and must-pass cases; every finding becomes a permanent test.

Try it#

  1. Write five adversarial cases for a system you know: one per category of your choice, each with a code-checkable success condition.
  2. Search a repository you maintain for places where model output reaches a shell, a query or an HTML template. How many validate it first?
  3. Define the must-pass list for your release gate.

Check yourself#

  1. Why should adversarial tests grade outcomes rather than the model’s wording?
  2. What is the difference between model-level and system-level attack success rate?
  3. Why does a model upgrade need to pass the security gate again?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom