Pidoku

Threat Modeling an AI System

Foundations 45 min Difficulty 2/5 Lesson 02 of 03

Prerequisites Why AI Security Is Different

The idea in one minute#

A threat model is a structured answer to four questions: what are we building, what can go wrong, what will we do about it, and did we do enough? For AI systems the method needs one adaptation. Classic models follow data flows; an AI model must also follow influence flows — which text can steer which model, and what that model can then do. Draw the system, mark who controls each piece of text entering a context, mark every effect a model can cause, and connect them. The connections are your threats.

A picture#

flowchart LR
  subgraph ACTORS["Who might attack"]
    direction TB
    X1[":i-user: <b>A user</b><br/><small>direct prompts</small>"]
    X2[":i-globe: <b>An outsider</b><br/><small>content the system reads</small>"]
    X3[":i-package: <b>A supplier</b><br/><small>model, tool, dependency</small>"]
    X4[":i-building-2: <b>An insider</b><br/><small>data, prompts, config</small>"]
  end
  subgraph ASSETS["What they want"]
    direction TB
    A1[(":i-database: <b>Data</b><br/><small>customer, company</small>")]
    A2[":i-key-round: <b>Credentials</b>"]
    A3[":i-wrench: <b>Actions</b><br/><small>pay, send, delete, deploy</small>"]
    A4[":i-coins: <b>Compute</b><br/><small>your GPU and API bill</small>"]
    A5[":i-brain: <b>The model</b><br/><small>weights, behaviour, reputation</small>"]
  end
  X1 --> SYS[":i-bot: <b>The AI system</b><br/><small>contexts, tools, stores</small>"]
  X2 --> SYS
  X3 --> SYS
  X4 --> SYS
  SYS --> A1
  SYS --> A2
  SYS --> A3
  SYS --> A4
  SYS --> A5
  class X1 queue
  class X2,X3 warn
  class X4 neutral
  class SYS compute
  class A1,A2,A3,A4,A5 memory

How it really works#

Step 1 — draw the system#

One diagram with every component that touches a model: where prompts are assembled, each model call, each tool, each store (documents, vectors, memory, logs), each external service. Draw trust boundaries — lines where control changes hands: the internet to your gateway, your service to a third-party model API, your agent to a third-party MCP server, tenant to tenant.

Step 2 — list the assets#

AssetExamples
DataCustomer records, source code, documents in the retrieval corpus, conversation history
CredentialsModel API keys, tool tokens, cloud roles reachable from the agent
ActionsAnything a tool can do that costs money, sends information or cannot be undone
ComputeGPU time and API spend an attacker can consume
The model itselfProprietary weights, system prompts, and the trust users place in its output
AvailabilityThe service staying up and affordable

Step 3 — list the actors#

ActorCan doTypical goal
Malicious userType anything; upload filesJailbreak, extract the system prompt, abuse free compute, reach other users’ data
Outsider with no accountPublish a web page, send an email, open a ticket, file a pull requestIndirect injection: act with a victim’s authority
SupplierPublish a model, dataset, package or MCP server you installBackdoor, credential theft, persistence
InsiderEdit prompts, corpora, training data, configurationSabotage, data theft
Another tenantShare your infrastructureCross-tenant leakage through caches, indexes, GPUs
Honest mistake—Not an attacker, and the most frequent cause of excessive agency doing damage

The outsider row is the one that classic models miss. They never log in, yet their text reaches your model.

Step 4 — trace influence#

For each model context in the diagram, fill in a row:

context: support agent, ticket-handling session

text sources          controlled by         trust
────────────          ─────────────         ─────
system prompt         us                    trusted
ticket body           the customer          UNTRUSTED
order lookup result   our database          trusted data, may contain customer-written fields → UNTRUSTED
knowledge base        support staff         trusted (review who can edit)
memory                previous sessions     as trusted as what could write it

effects available     class
─────────────────     ─────
lookup_order          read (private data)
issue_refund          write, irreversible, money
send_reply            send, external
update_memory         write, persistent

Then ask of every untrusted source and every effect: can this text cause that effect? In the example, customer-written ticket text can reach issue_refund and send_reply in the same context. Those two lines are the threats to mitigate first.

Step 5 — enumerate what can go wrong#

STRIDE, the classic checklist, still works when each letter is read with the model in mind:

STRIDEIn an AI system
SpoofingAn agent or MCP server claiming an identity it does not have; content impersonating the system or the user inside the context
TamperingPoisoned training data, corpus, memory or tool descriptions; modified weights
RepudiationNo record of which agent did what for whom, or of what was in its context
Information disclosureExfiltration through tool calls or rendered links; system-prompt and cross-tenant leakage; training-data extraction
Denial of serviceUnbounded token consumption; agent loops; denial of wallet
Elevation of privilegeInjection that makes the model use its authority for the attacker; excessive agency

Step 6 — rate and decide#

Rate each threat on impact (what is lost) and likelihood (how reachable is it: can any anonymous outsider trigger it, or only an authenticated insider?). For AI threats, add a third consideration that changes priorities sharply:

Is the mitigation deterministic or probabilistic?

A high-impact threat mitigated only by “the model was told not to” or “a classifier usually catches it” is not mitigated. It needs a control from this list:

  • Remove the capability (the agent has no such tool).
  • Remove the data (the agent cannot read it).
  • Remove the path (the untrusted source and the effect are in different contexts).
  • Bound it (scoped credentials, limits, sandbox, egress allowlist).
  • Gate it (human approval on the exact action).
  • Make it reversible (draft, branch, soft delete).

Step 7 — write it down and keep it alive#

A threat model is a short document: the diagram, the influence table per context, the ranked threats with the control for each, and the accepted residual risks with an owner. It is revisited whenever the system changes in one of these ways:

  • a new tool or MCP server,
  • a new source of context text,
  • a new store the model can write,
  • a change in whose authority the agent uses,
  • a move toward more autonomy.

Adding a tool is the commonest way a safe system becomes an unsafe one. A harmless summariser plus a helpful “send this to my colleague” button is a data-exfiltration path.

Code#

Influence tracing as a program: given the sources and effects of each context, report which contexts hold the dangerous combination.

Go
// trifecta.go — find agent contexts that combine private data, untrusted content and a way out.
package main

import "fmt"

type agentContext struct {
	name      string
	sources   map[string]bool // source → is it attacker-controllable?
	effects   map[string]string
	approvals map[string]bool // effect → needs human approval
}

func (c agentContext) analyse() {
	untrusted, private, outbound := []string{}, []string{}, []string{}
	for s, bad := range c.sources {
		if bad {
			untrusted = append(untrusted, s)
		}
	}
	for e, class := range c.effects {
		switch class {
		case "read-private":
			private = append(private, e)
		case "send", "write-irreversible":
			if !c.approvals[e] {
				outbound = append(outbound, e)
			}
		}
	}
	legs := 0
	for _, l := range [][]string{untrusted, private, outbound} {
		if len(l) > 0 {
			legs++
		}
	}
	verdict := "ok"
	if legs == 3 {
		verdict = "UNSAFE: an attacker's text can reach private data and an ungated way out"
	}
	fmt.Printf("%-22s legs=%d  %s\n", c.name, legs, verdict)
	if legs == 3 {
		fmt.Printf("    untrusted: %v\n    private:   %v\n    ungated:   %v\n", untrusted, private, outbound)
	}
}

func main() {
	contexts := []agentContext{
		{"web summariser",
			map[string]bool{"system prompt": false, "web page": true},
			map[string]string{}, nil},
		{"support agent",
			map[string]bool{"system prompt": false, "ticket body": true, "knowledge base": false},
			map[string]string{"lookup_order": "read-private", "send_reply": "send", "issue_refund": "write-irreversible"},
			nil},
		{"support agent, gated",
			map[string]bool{"system prompt": false, "ticket body": true, "knowledge base": false},
			map[string]string{"lookup_order": "read-private", "send_reply": "send", "issue_refund": "write-irreversible"},
			map[string]bool{"send_reply": true, "issue_refund": true}},
	}
	for _, c := range contexts {
		c.analyse()
	}
}

The same agent is unsafe or acceptable depending only on whether its outbound effects are gated. That is a property of the architecture; no prompt was involved.

Remember this#

  • Follow influence, not just data: which text can steer which model, and what can that model do.
  • The outsider who never logs in is the actor classic models miss.
  • A threat mitigated only by instructions or a classifier is not mitigated.
  • The fixes are architectural: remove capability, data or path; bound; gate; make reversible.
  • Redo the model whenever a tool, a source, a writable store or the agent’s authority changes.

Try it#

  1. Run trifecta.go. Add a context for a coding agent with a sandbox and no network; what does it report, and what single change would make it unsafe?
  2. Build the influence table for one model context in a system you know.
  3. Take one threat from your table and choose a deterministic control for it from the list.

Check yourself#

  1. What is an influence flow, and how does it differ from a data flow?
  2. Why does adding a tool require revisiting the threat model?
  3. Give an example of each kind of deterministic mitigation.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom