The idea in one minute#
A threat model is a structured answer to four questions: what are we building, what can go wrong, what will we do about it, and did we do enough? For AI systems the method needs one adaptation. Classic models follow data flows; an AI model must also follow influence flows — which text can steer which model, and what that model can then do. Draw the system, mark who controls each piece of text entering a context, mark every effect a model can cause, and connect them. The connections are your threats.
A picture#
flowchart LR
subgraph ACTORS["Who might attack"]
direction TB
X1[":i-user: <b>A user</b><br/><small>direct prompts</small>"]
X2[":i-globe: <b>An outsider</b><br/><small>content the system reads</small>"]
X3[":i-package: <b>A supplier</b><br/><small>model, tool, dependency</small>"]
X4[":i-building-2: <b>An insider</b><br/><small>data, prompts, config</small>"]
end
subgraph ASSETS["What they want"]
direction TB
A1[(":i-database: <b>Data</b><br/><small>customer, company</small>")]
A2[":i-key-round: <b>Credentials</b>"]
A3[":i-wrench: <b>Actions</b><br/><small>pay, send, delete, deploy</small>"]
A4[":i-coins: <b>Compute</b><br/><small>your GPU and API bill</small>"]
A5[":i-brain: <b>The model</b><br/><small>weights, behaviour, reputation</small>"]
end
X1 --> SYS[":i-bot: <b>The AI system</b><br/><small>contexts, tools, stores</small>"]
X2 --> SYS
X3 --> SYS
X4 --> SYS
SYS --> A1
SYS --> A2
SYS --> A3
SYS --> A4
SYS --> A5
class X1 queue
class X2,X3 warn
class X4 neutral
class SYS compute
class A1,A2,A3,A4,A5 memoryHow it really works#
Step 1 — draw the system#
One diagram with every component that touches a model: where prompts are assembled, each model call, each tool, each store (documents, vectors, memory, logs), each external service. Draw trust boundaries — lines where control changes hands: the internet to your gateway, your service to a third-party model API, your agent to a third-party MCP server, tenant to tenant.
Step 2 — list the assets#
| Asset | Examples |
|---|---|
| Data | Customer records, source code, documents in the retrieval corpus, conversation history |
| Credentials | Model API keys, tool tokens, cloud roles reachable from the agent |
| Actions | Anything a tool can do that costs money, sends information or cannot be undone |
| Compute | GPU time and API spend an attacker can consume |
| The model itself | Proprietary weights, system prompts, and the trust users place in its output |
| Availability | The service staying up and affordable |
Step 3 — list the actors#
| Actor | Can do | Typical goal |
|---|---|---|
| Malicious user | Type anything; upload files | Jailbreak, extract the system prompt, abuse free compute, reach other users’ data |
| Outsider with no account | Publish a web page, send an email, open a ticket, file a pull request | Indirect injection: act with a victim’s authority |
| Supplier | Publish a model, dataset, package or MCP server you install | Backdoor, credential theft, persistence |
| Insider | Edit prompts, corpora, training data, configuration | Sabotage, data theft |
| Another tenant | Share your infrastructure | Cross-tenant leakage through caches, indexes, GPUs |
| Honest mistake | — | Not an attacker, and the most frequent cause of excessive agency doing damage |
The outsider row is the one that classic models miss. They never log in, yet their text reaches your model.
Step 4 — trace influence#
For each model context in the diagram, fill in a row:
context: support agent, ticket-handling session
text sources controlled by trust
──────────── ───────────── ─────
system prompt us trusted
ticket body the customer UNTRUSTED
order lookup result our database trusted data, may contain customer-written fields → UNTRUSTED
knowledge base support staff trusted (review who can edit)
memory previous sessions as trusted as what could write it
effects available class
───────────────── ─────
lookup_order read (private data)
issue_refund write, irreversible, money
send_reply send, external
update_memory write, persistentThen ask of every untrusted source and every effect: can this text cause that effect? In
the example, customer-written ticket text can reach issue_refund and send_reply in the
same context. Those two lines are the threats to mitigate first.
Step 5 — enumerate what can go wrong#
STRIDE, the classic checklist, still works when each letter is read with the model in mind:
| STRIDE | In an AI system |
|---|---|
| Spoofing | An agent or MCP server claiming an identity it does not have; content impersonating the system or the user inside the context |
| Tampering | Poisoned training data, corpus, memory or tool descriptions; modified weights |
| Repudiation | No record of which agent did what for whom, or of what was in its context |
| Information disclosure | Exfiltration through tool calls or rendered links; system-prompt and cross-tenant leakage; training-data extraction |
| Denial of service | Unbounded token consumption; agent loops; denial of wallet |
| Elevation of privilege | Injection that makes the model use its authority for the attacker; excessive agency |
Step 6 — rate and decide#
Rate each threat on impact (what is lost) and likelihood (how reachable is it: can any anonymous outsider trigger it, or only an authenticated insider?). For AI threats, add a third consideration that changes priorities sharply:
Is the mitigation deterministic or probabilistic?
A high-impact threat mitigated only by “the model was told not to” or “a classifier usually catches it” is not mitigated. It needs a control from this list:
- Remove the capability (the agent has no such tool).
- Remove the data (the agent cannot read it).
- Remove the path (the untrusted source and the effect are in different contexts).
- Bound it (scoped credentials, limits, sandbox, egress allowlist).
- Gate it (human approval on the exact action).
- Make it reversible (draft, branch, soft delete).
Step 7 — write it down and keep it alive#
A threat model is a short document: the diagram, the influence table per context, the ranked threats with the control for each, and the accepted residual risks with an owner. It is revisited whenever the system changes in one of these ways:
- a new tool or MCP server,
- a new source of context text,
- a new store the model can write,
- a change in whose authority the agent uses,
- a move toward more autonomy.
Adding a tool is the commonest way a safe system becomes an unsafe one. A harmless summariser plus a helpful “send this to my colleague” button is a data-exfiltration path.
Code#
Influence tracing as a program: given the sources and effects of each context, report which contexts hold the dangerous combination.
// trifecta.go — find agent contexts that combine private data, untrusted content and a way out.
package main
import "fmt"
type agentContext struct {
name string
sources map[string]bool // source → is it attacker-controllable?
effects map[string]string
approvals map[string]bool // effect → needs human approval
}
func (c agentContext) analyse() {
untrusted, private, outbound := []string{}, []string{}, []string{}
for s, bad := range c.sources {
if bad {
untrusted = append(untrusted, s)
}
}
for e, class := range c.effects {
switch class {
case "read-private":
private = append(private, e)
case "send", "write-irreversible":
if !c.approvals[e] {
outbound = append(outbound, e)
}
}
}
legs := 0
for _, l := range [][]string{untrusted, private, outbound} {
if len(l) > 0 {
legs++
}
}
verdict := "ok"
if legs == 3 {
verdict = "UNSAFE: an attacker's text can reach private data and an ungated way out"
}
fmt.Printf("%-22s legs=%d %s\n", c.name, legs, verdict)
if legs == 3 {
fmt.Printf(" untrusted: %v\n private: %v\n ungated: %v\n", untrusted, private, outbound)
}
}
func main() {
contexts := []agentContext{
{"web summariser",
map[string]bool{"system prompt": false, "web page": true},
map[string]string{}, nil},
{"support agent",
map[string]bool{"system prompt": false, "ticket body": true, "knowledge base": false},
map[string]string{"lookup_order": "read-private", "send_reply": "send", "issue_refund": "write-irreversible"},
nil},
{"support agent, gated",
map[string]bool{"system prompt": false, "ticket body": true, "knowledge base": false},
map[string]string{"lookup_order": "read-private", "send_reply": "send", "issue_refund": "write-irreversible"},
map[string]bool{"send_reply": true, "issue_refund": true}},
}
for _, c := range contexts {
c.analyse()
}
}The same agent is unsafe or acceptable depending only on whether its outbound effects are gated. That is a property of the architecture; no prompt was involved.
Remember this#
- Follow influence, not just data: which text can steer which model, and what can that model do.
- The outsider who never logs in is the actor classic models miss.
- A threat mitigated only by instructions or a classifier is not mitigated.
- The fixes are architectural: remove capability, data or path; bound; gate; make reversible.
- Redo the model whenever a tool, a source, a writable store or the agent’s authority changes.
Try it#
- Run
trifecta.go. Add a context for a coding agent with a sandbox and no network; what does it report, and what single change would make it unsafe? - Build the influence table for one model context in a system you know.
- Take one threat from your table and choose a deterministic control for it from the list.
Check yourself#
- What is an influence flow, and how does it differ from a data flow?
- Why does adding a tool require revisiting the threat model?
- Give an example of each kind of deterministic mitigation.