The idea in one minute#
Security controls for an AI application attach to seven stages: design, data, model, build, deploy, runtime, operate. Each attack from the previous topic is cheapest to stop at one particular stage — a poisoned model at the registry gate, an over-privileged tool at design, an injection’s consequences at runtime policy. This lesson is the map: which control lives where, what evidence it leaves, and who owns it. The six lessons that follow each take one stage.
The map also makes one thing visible: runtime guardrails, where most attention goes, are one column out of seven.
A picture#
flowchart LR S1[":i-list-checks: <b>1 Design</b><br/><small>threat model,<br/>least agency</small>"] --> S2[":i-database: <b>2 Data</b><br/><small>provenance, privacy,<br/>access</small>"] S2 --> S3[":huggingface: <b>3 Model</b><br/><small>scan, sign,<br/>registry</small>"] S3 --> S4[":github: <b>4 Build</b><br/><small>eval gates,<br/>red team in CI</small>"] S4 --> S5[":kubernetes: <b>5 Deploy</b><br/><small>isolation, secrets,<br/>network</small>"] S5 --> S6[":i-shield-check: <b>6 Runtime</b><br/><small>guardrails, policy,<br/>limits</small>"] S6 --> S7[":i-radar: <b>7 Operate</b><br/><small>detect, respond,<br/>learn</small>"] S7 -.->|"every incident becomes a test and a threat-model update"| S1 class S1 neutral class S2,S3 memory class S4 queue class S5 compute class S6 warn class S7 io
How it really works#
The map#
| Stage | Main threats | Key controls | Evidence it leaves |
|---|---|---|---|
| 1 Design | Excessive agency; the trifecta; missing trust boundaries | Threat model; least-privilege tool set; choice of autonomy level; data-flow and influence-flow diagram | A reviewed threat model, updated on every new tool or source |
| 2 Data | Poisoning; privacy leakage; unauthorised use | Provenance records; access control on corpora; PII handling; ingestion scanning; dataset versioning | Dataset manifests, lineage, consent and licence records |
| 3 Model | Malicious model files; backdoors; tampering | Weights-only formats; scanning; signatures; internal registry; AI bill of materials | Signed artifacts, scan reports, AI-BOM |
| 4 Build | Regressions in safety behaviour; vulnerable dependencies; secrets in prompts | Prompts and policies as code; security evaluation suite; automated red teaming; dependency and secret scanning | CI results per commit; release gate decisions |
| 5 Deploy | Credential theft; cross-tenant leakage; exposed endpoints; lateral movement | Workload identity; secret management; network policy; sandboxing; GPU and cache isolation | Infrastructure as code; policy reports; attestation |
| 6 Runtime | Injection; exfiltration; jailbreak; denial of wallet | Input and output screening; tool-call policy; output sanitising; token and spend limits; approvals | Policy decisions and guardrail verdicts per request |
| 7 Operate | Undetected compromise; slow response; repeat incidents | Audit trail; behavioural detection; kill switches; incident runbooks; post-incident tests | Alerts, incident records, new test cases |
Shift left, and also right#
“Shift left” — catch problems early — holds here: removing a dangerous tool at design costs a meeting, and discovering it in an incident costs a breach. But AI systems also need a strong right side, for two reasons:
- The model’s behaviour cannot be fully verified before release. Testing samples a space that cannot be enumerated, so some failures will only appear in production.
- The inputs change after release. New documents enter the corpus, tools update, users find new phrasings. A system that was safe at launch drifts.
So the lifecycle is a loop: operate feeds design.
Deterministic and probabilistic, by stage#
| Stage | Deterministic controls | Probabilistic controls |
|---|---|---|
| Design | Capability removal; privilege scoping | — |
| Data | Access control; allowlisted sources; hashes | Poison and PII detectors |
| Model | Format restrictions; signatures; registry policy | Backdoor scanning; behavioural evaluation |
| Build | Gates on must-pass tests; secret scanning | Red-team success rates |
| Deploy | Network policy; isolation; identity | — |
| Runtime | Tool policy; limits; sanitisers; approvals | Classifiers; judge models |
| Operate | Kill switches; revocation | Anomaly detection |
The left column is what a security review can sign off. The right column is what makes attacks rarer and incidents more visible. A design leaning mainly on the right column for a high-impact threat needs rework.
Who owns what#
AI security falls between teams unless ownership is written down.
| Owner | Typically owns |
|---|---|
| Product / application team | Threat model, tool set, prompts, evaluation suite, approval design |
| Data team | Dataset provenance, corpus access control, PII handling |
| ML or platform team | Model registry, serving, gateway, sandbox and identity infrastructure |
| Security team | Standards, review, red teaming, detection, incident response |
| Legal and compliance | Regulatory classification, data-use terms, records |
Two things that most often have no owner: the MCP server and tool catalogue (who approves a new one?) and the memory store (who can inspect and clean it?). Assign both.
A minimum bar#
If you can only do ten things, do these, in this order:
- Put every model call behind a gateway with authentication, token limits and spend budgets.
- Keep credentials out of prompts, code and clients; issue short-lived, scoped ones.
- Write the threat model; remove tools the task does not need.
- Enforce data permissions before the context, in retrieval and in tools.
- Gate irreversible and outbound actions with policy in code, and approval where needed.
- Sanitise rendered output; restrict egress from anything that executes code.
- Load models only from your registry, in a weights-only format.
- Approve and pin every tool and MCP server.
- Log every action with user, agent, arguments and context sources.
- Add adversarial cases to the evaluation suite and run them on every change.
Guardrail classifiers are useful and are not on this list: they come after the controls that bound damage.
Maturity, in three steps#
| Level | Looks like |
|---|---|
| Baseline | The ten items above; a named owner; an incident runbook |
| Managed | Automated red teaming in CI; signed models and AI-BOMs; per-task credentials; behavioural detection; regular exercises |
| Advanced | Architectural isolation of untrusted content; information-flow controls; confidential computing for sensitive tenants; continuous adversarial evaluation in production |
Remember this#
- Seven stages: design, data, model, build, deploy, runtime, operate.
- Each attack has a cheapest stage to stop it; runtime guardrails are one stage of seven.
- The lifecycle is a loop, because behaviour cannot be fully verified before release.
- Sign off on deterministic controls; use probabilistic ones to lower rates and raise visibility.
- Assign owners — especially for the tool catalogue and the memory store.
Try it#
- Score a system you know against the ten-item minimum bar.
- For the three highest threats in your threat model, name the stage where each is cheapest to stop and the control you would place there.
- Write down who owns tool approval and memory hygiene in your organisation. If the answer is “nobody”, propose an owner.
Check yourself#
- Why does an AI system need strong controls on the “right” side of the lifecycle as well as the left?
- Which stage stops a malicious model file most cheaply?
- Why are guardrail classifiers not in the ten-item minimum?