Pidoku

Trust Boundaries in an AI System

Advanced 45 min Difficulty 3/5 Lesson 01 of 02

Prerequisites Anatomy of an Agent, Tools and Protocols

The idea in one minute#

In ordinary software, code is trusted and data is not, and the two travel in separate channels. In an AI system they travel in one: everything in the context window — your instructions, the user’s message, a retrieved document, a tool result — is text the model reads, and any of it can steer what the model does next. So the security question for an AI design is not “is the input sanitised?” It is: what can reach the context, what can the model’s output cause, and what stands between the two?

Draw those three things on the design and most of the security work is visible.

A picture#

flowchart LR
  subgraph SRC["Sources: who controls this text?"]
    direction LR
    S1[":i-code: <b>System prompt</b><br/><small>you: trusted</small>"]
    S2[":i-user: <b>User message</b><br/><small>the user: partly trusted</small>"]
    S3[":i-globe: <b>Web pages, email, files</b><br/><small>anyone: untrusted</small>"]
    S4[":modelcontextprotocol: <b>Tool results</b><br/><small>the tool's data: untrusted</small>"]
    S5[(":i-database: <b>Retrieved docs, memory</b><br/><small>whoever wrote them</small>")]
  end
  SRC --> CTX[":i-layers: <b>Context window</b><br/><small>one channel, no separation</small>"]
  CTX --> M[":i-brain: <b>Model</b>"]
  M --> POL[":i-shield-check: <b>Policy layer</b><br/><small>deterministic code</small>"]
  subgraph SINK["Sinks: what can output cause?"]
    direction LR
    K1[":i-eye: <b>Shown to the user</b><br/><small>links, images, markup</small>"]
    K2[":i-wrench: <b>Tool calls</b><br/><small>read, write, send, pay</small>"]
    K3[":i-terminal: <b>Code execution</b>"]
    K4[(":i-archive: <b>Memory writes</b>")]
  end
  POL --> SINK
  class S1 neutral
  class S2 queue
  class S3,S4,S5 warn
  class CTX io
  class M compute
  class POL queue
  class K1,K2,K3,K4 warn

How it really works#

Step 1 — list the sources and who controls them#

SourceControlled byTrust
System prompt, tool definitions you wroteYouTrusted
The user’s messageThe authenticated userTrusted to act for themselves, not for others
Uploaded filesWhoever created the file — often not the userUntrusted
Web pages, search resultsAnyone on the internetUntrusted
Email, tickets, chat messages, calendar invitesAnyone who can send oneUntrusted
Tool and API resultsWhoever can write to the underlying dataUntrusted
Retrieved documentsEvery author in the corpusAs trusted as the least trusted author
MemoryWhatever was allowed to write itAs trusted as its write path
Third-party MCP server descriptionsThe server’s publisherA supplier you must vet
Another agent’s outputThat agent’s operator, and its inputsUntrusted

The uncomfortable row is tool results. A “read the ticket” tool returns text written by a customer; a “fetch the page” tool returns text written by anyone. An agent with tools is an agent that reads untrusted text.

Step 2 — list the sinks and what each can cause#

SinkWorst case
Text rendered to the userA link or image whose URL carries stolen data out; misleading content
A read toolAccess to data beyond what the task needs
A write toolChanged or destroyed data
A communication tool: email, chat, HTTP requestData sent to an attacker
Code executionAnything the sandbox permits
A memory writeA planted instruction that fires in later sessions
A call to another agentThe problem, forwarded

Step 3 — find the dangerous paths#

A design is at risk wherever one agent context combines three things:

   A   access to private data
   B   exposure to untrusted content
   C   a way to send data out or change state

With all three, an attacker who controls the untrusted content can instruct the agent to take the private data and send it out. This combination is widely called the lethal trifecta, and a practical design rule follows from it — often stated as the rule of two: within one session, an agent may have at most two of the three without a human approving the action.

ConfigurationABCVerdict
Summarise public web pages–✓–Safe: nothing to steal, no way out
Internal assistant over company documents, no outbound tools✓––Safe while every author is trusted
Coding agent in a sandbox with no network✓✓–Safe: data cannot leave
Email assistant that reads mail and can send mail✓✓✓Unsafe without approval on sending
Browser agent logged into the user’s accounts✓✓✓Unsafe by construction; needs strong gating

C is broader than it looks. Rendering a markdown image, following a link, making a DNS lookup or writing to a shared document are all ways to send data out.

Step 4 — place the controls#

Controls belong at boundaries, and they are of two kinds. Probabilistic controls — classifiers, guardrail models, instructions in the prompt — reduce how often attacks succeed. Deterministic controls — permissions, isolation, allowlists, approvals — bound what a successful attack can do. You need both, and only the second kind can be relied on.

BoundaryDeterministic controlsProbabilistic controls
Source → contextProvenance labels; only allow trusted writers for memory and instructionsInjection classifiers on retrieved content
Model → sinkPolicy layer: schema validation, per-tool authorization, argument allowlistsOutput classifiers
Around the agentLeast-privilege identity, scoped short-lived tokensAnomaly detection on behaviour
Around executionSandbox, egress allowlist, no secrets inside—
Before irreversible effectsHuman approval bound to the exact action—
Around the user’s viewSanitise rendered output; block untrusted image and link hosts—

important

A system prompt that says “never reveal customer data” is a probabilistic control. An agent identity that has no permission to read other customers’ data is a deterministic one. Design so that the first can fail without consequence.

Step 5 — bound the blast radius#

Assume the model will, at some point, be successfully manipulated, and ask what the damage is. The answer should be small because of the design:

  • Act as the user, not as the service. The agent’s tools run with the requesting user’s permissions, so it can never reach more than that user could.
  • Scope to the task. A token for one repository, one ticket, one mailbox — not all of them.
  • Short lifetimes. Credentials and sandboxes die with the task.
  • Separate the reader from the actor. The component that processes untrusted content has no powerful tools; the component with powerful tools sees only structured, validated data.
  • Make it reversible. Drafts instead of sends, branches instead of pushes to main, soft deletes.
  • Log everything. Each action is recorded with the agent, the user it acted for, and the inputs that led to it.

The threat-model worksheet#

Fill this in for every AI design. It takes an hour and finds most problems.

1. Sources     every kind of text entering any context, and who controls it
2. Sinks       every effect model output can cause, classified read / write / send / execute
3. Paths       for each untrusted source, which sinks are reachable in the same context
4. Trifecta    for each agent context: private data? untrusted content? outbound channel?
5. Controls    for each path: the deterministic control, and what happens when the
               probabilistic ones fail
6. Identity    whose authority each tool call carries, and how it is scoped
7. Blast       the worst outcome of a fully hijacked agent, in one sentence
8. Detection   how you would know it had happened

Remember this#

  • One channel carries instructions and data; anything in the context can steer the model.
  • Tool results and retrieved documents are untrusted input.
  • The lethal trifecta: private data + untrusted content + an outbound channel. Allow at most two without a human.
  • Probabilistic controls lower the odds; deterministic controls bound the damage.
  • Design for the case where the model has been manipulated.

Try it#

  1. Fill in the worksheet for an assistant that reads a shared inbox and drafts replies.
  2. For a product you use that has an AI assistant, find its sources and sinks. Does any context hold all three legs of the trifecta?
  3. Take one unsafe row of the configuration table and redesign it to remove one leg.

Check yourself#

  1. Why is a tool result untrusted even when the tool itself is yours?
  2. Give three outbound channels that are not obviously “sending data”.
  3. What is the difference between a probabilistic and a deterministic control, and why does it matter which one a design relies on?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom