The idea in one minute#
In ordinary software, code is trusted and data is not, and the two travel in separate channels. In an AI system they travel in one: everything in the context window — your instructions, the user’s message, a retrieved document, a tool result — is text the model reads, and any of it can steer what the model does next. So the security question for an AI design is not “is the input sanitised?” It is: what can reach the context, what can the model’s output cause, and what stands between the two?
Draw those three things on the design and most of the security work is visible.
A picture#
flowchart LR
subgraph SRC["Sources: who controls this text?"]
direction LR
S1[":i-code: <b>System prompt</b><br/><small>you: trusted</small>"]
S2[":i-user: <b>User message</b><br/><small>the user: partly trusted</small>"]
S3[":i-globe: <b>Web pages, email, files</b><br/><small>anyone: untrusted</small>"]
S4[":modelcontextprotocol: <b>Tool results</b><br/><small>the tool's data: untrusted</small>"]
S5[(":i-database: <b>Retrieved docs, memory</b><br/><small>whoever wrote them</small>")]
end
SRC --> CTX[":i-layers: <b>Context window</b><br/><small>one channel, no separation</small>"]
CTX --> M[":i-brain: <b>Model</b>"]
M --> POL[":i-shield-check: <b>Policy layer</b><br/><small>deterministic code</small>"]
subgraph SINK["Sinks: what can output cause?"]
direction LR
K1[":i-eye: <b>Shown to the user</b><br/><small>links, images, markup</small>"]
K2[":i-wrench: <b>Tool calls</b><br/><small>read, write, send, pay</small>"]
K3[":i-terminal: <b>Code execution</b>"]
K4[(":i-archive: <b>Memory writes</b>")]
end
POL --> SINK
class S1 neutral
class S2 queue
class S3,S4,S5 warn
class CTX io
class M compute
class POL queue
class K1,K2,K3,K4 warnHow it really works#
Step 1 — list the sources and who controls them#
| Source | Controlled by | Trust |
|---|---|---|
| System prompt, tool definitions you wrote | You | Trusted |
| The user’s message | The authenticated user | Trusted to act for themselves, not for others |
| Uploaded files | Whoever created the file — often not the user | Untrusted |
| Web pages, search results | Anyone on the internet | Untrusted |
| Email, tickets, chat messages, calendar invites | Anyone who can send one | Untrusted |
| Tool and API results | Whoever can write to the underlying data | Untrusted |
| Retrieved documents | Every author in the corpus | As trusted as the least trusted author |
| Memory | Whatever was allowed to write it | As trusted as its write path |
| Third-party MCP server descriptions | The server’s publisher | A supplier you must vet |
| Another agent’s output | That agent’s operator, and its inputs | Untrusted |
The uncomfortable row is tool results. A “read the ticket” tool returns text written by a customer; a “fetch the page” tool returns text written by anyone. An agent with tools is an agent that reads untrusted text.
Step 2 — list the sinks and what each can cause#
| Sink | Worst case |
|---|---|
| Text rendered to the user | A link or image whose URL carries stolen data out; misleading content |
| A read tool | Access to data beyond what the task needs |
| A write tool | Changed or destroyed data |
| A communication tool: email, chat, HTTP request | Data sent to an attacker |
| Code execution | Anything the sandbox permits |
| A memory write | A planted instruction that fires in later sessions |
| A call to another agent | The problem, forwarded |
Step 3 — find the dangerous paths#
A design is at risk wherever one agent context combines three things:
A access to private data
B exposure to untrusted content
C a way to send data out or change stateWith all three, an attacker who controls the untrusted content can instruct the agent to take the private data and send it out. This combination is widely called the lethal trifecta, and a practical design rule follows from it — often stated as the rule of two: within one session, an agent may have at most two of the three without a human approving the action.
| Configuration | A | B | C | Verdict |
|---|---|---|---|---|
| Summarise public web pages | – | ✓ | – | Safe: nothing to steal, no way out |
| Internal assistant over company documents, no outbound tools | ✓ | – | – | Safe while every author is trusted |
| Coding agent in a sandbox with no network | ✓ | ✓ | – | Safe: data cannot leave |
| Email assistant that reads mail and can send mail | ✓ | ✓ | ✓ | Unsafe without approval on sending |
| Browser agent logged into the user’s accounts | ✓ | ✓ | ✓ | Unsafe by construction; needs strong gating |
C is broader than it looks. Rendering a markdown image, following a link, making a DNS lookup or writing to a shared document are all ways to send data out.
Step 4 — place the controls#
Controls belong at boundaries, and they are of two kinds. Probabilistic controls — classifiers, guardrail models, instructions in the prompt — reduce how often attacks succeed. Deterministic controls — permissions, isolation, allowlists, approvals — bound what a successful attack can do. You need both, and only the second kind can be relied on.
| Boundary | Deterministic controls | Probabilistic controls |
|---|---|---|
| Source → context | Provenance labels; only allow trusted writers for memory and instructions | Injection classifiers on retrieved content |
| Model → sink | Policy layer: schema validation, per-tool authorization, argument allowlists | Output classifiers |
| Around the agent | Least-privilege identity, scoped short-lived tokens | Anomaly detection on behaviour |
| Around execution | Sandbox, egress allowlist, no secrets inside | — |
| Before irreversible effects | Human approval bound to the exact action | — |
| Around the user’s view | Sanitise rendered output; block untrusted image and link hosts | — |
important
A system prompt that says “never reveal customer data” is a probabilistic control. An agent identity that has no permission to read other customers’ data is a deterministic one. Design so that the first can fail without consequence.
Step 5 — bound the blast radius#
Assume the model will, at some point, be successfully manipulated, and ask what the damage is. The answer should be small because of the design:
- Act as the user, not as the service. The agent’s tools run with the requesting user’s permissions, so it can never reach more than that user could.
- Scope to the task. A token for one repository, one ticket, one mailbox — not all of them.
- Short lifetimes. Credentials and sandboxes die with the task.
- Separate the reader from the actor. The component that processes untrusted content has no powerful tools; the component with powerful tools sees only structured, validated data.
- Make it reversible. Drafts instead of sends, branches instead of pushes to main, soft deletes.
- Log everything. Each action is recorded with the agent, the user it acted for, and the inputs that led to it.
The threat-model worksheet#
Fill this in for every AI design. It takes an hour and finds most problems.
1. Sources every kind of text entering any context, and who controls it
2. Sinks every effect model output can cause, classified read / write / send / execute
3. Paths for each untrusted source, which sinks are reachable in the same context
4. Trifecta for each agent context: private data? untrusted content? outbound channel?
5. Controls for each path: the deterministic control, and what happens when the
probabilistic ones fail
6. Identity whose authority each tool call carries, and how it is scoped
7. Blast the worst outcome of a fully hijacked agent, in one sentence
8. Detection how you would know it had happenedRemember this#
- One channel carries instructions and data; anything in the context can steer the model.
- Tool results and retrieved documents are untrusted input.
- The lethal trifecta: private data + untrusted content + an outbound channel. Allow at most two without a human.
- Probabilistic controls lower the odds; deterministic controls bound the damage.
- Design for the case where the model has been manipulated.
Try it#
- Fill in the worksheet for an assistant that reads a shared inbox and drafts replies.
- For a product you use that has an AI assistant, find its sources and sinks. Does any context hold all three legs of the trifecta?
- Take one unsafe row of the configuration table and redesign it to remove one leg.
Check yourself#
- Why is a tool result untrusted even when the tool itself is yours?
- Give three outbound channels that are not obviously “sending data”.
- What is the difference between a probabilistic and a deterministic control, and why does it matter which one a design relies on?