Pidoku

Agent Threats

Advanced 50 min Difficulty 3/5 Lesson 01 of 05

Prerequisites How Attacks Work, Threat Modeling an AI System

The idea in one minute#

A chatbot that is fooled says something wrong. An agent that is fooled does something wrong, keeps doing it across many steps, may remember the instruction for next time, and may pass it to other agents. Four properties create the new risk: tools (effects), autonomy (no human between steps), memory (persistence) and composition (agents feeding agents). The OWASP Top 10 for Agentic Applications names the ten failure modes that follow. This lesson walks through them as one connected picture: how an attack enters, what it can reach, how it stays, and how it spreads.

A picture#

flowchart LR
  subgraph ENTER["Entry"]
    direction TB
    E1[":i-globe: <b>Untrusted content</b>"]
    E2[":modelcontextprotocol: <b>Poisoned tool or server</b>"]
    E3[":a2a: <b>Another agent's output</b>"]
  end
  ENTER --> HJ[":i-skull: <b>ASI01 Goal hijack</b>"]
  HJ --> TM[":i-wrench: <b>ASI02 Tool misuse</b>"]
  HJ --> ID[":i-key-round: <b>ASI03 Privilege abuse</b>"]
  HJ --> CE[":i-terminal: <b>ASI05 Code execution</b>"]
  HJ --> MP[(":i-archive: <b>ASI06 Memory poisoning</b>")]
  MP -->|"fires in later sessions"| HJ
  TM --> IA[":i-network: <b>ASI07 Inter-agent messages</b>"]
  IA --> CF[":i-zap: <b>ASI08 Cascading failure</b>"]
  TM --> HT[":i-hand: <b>ASI09 Human trust exploited</b>"]
  E2 --> SC[":i-package: <b>ASI04 Supply chain</b>"] --> HJ
  CF --> RG[":i-bot: <b>ASI10 Rogue agent</b>"]
  class E1,E2,E3 warn
  class HJ warn
  class TM,ID,CE io
  class MP memory
  class IA,CF queue
  class HT neutral
  class SC warn
  class RG warn

How it really works#

ASI01 — Agent goal hijack#

The agent’s objective is replaced by an attacker’s, through content it reads. This is indirect prompt injection with consequences: the agent now plans toward the new goal over many steps, choosing tools creatively to achieve it. Hijack can be total (“do this instead”) or partial and harder to notice (“also do this”, “prefer this vendor”, “skip the security check”).

Controls: treat everything read as untrusted; gate actions by session taint; isolate the reading of untrusted content from the holding of privileges — Architectural Defenses.

ASI02 — Tool misuse and exploitation#

Legitimate tools used in ways nobody intended. Nothing is “exploited” in the classic sense: the agent calls send_email exactly as designed, with the wrong content and recipient. Also here: unsafe chaining (read a secret with one tool, send it with another), abusing over-broad tools (run_sql instead of get_order), and exhausting a tool’s backend.

Controls: narrow, purpose-built tools; argument validation and allowlists; per-tool rate limits; least agency — the agent has only the tools this task needs.

ASI03 — Identity and privilege abuse#

The agent holds more authority than the request warrants, or the wrong party’s authority. Typical flaws: a single service account used for every user; a token minted for one tool accepted by another; privileges inherited by sub-agents that should not have them; an agent with a user’s full delegated access when the task needed one mailbox folder.

Controls: the agent acts on behalf of the user with scoped, short-lived, audience-bound tokens — Identity and Authorization.

ASI04 — Agentic supply chain vulnerabilities#

Compromised components that the agent loads at run time: MCP servers, tools, skills, plugins, prompt packs, models, other agents’ cards. Dynamic discovery makes it worse — an agent that can find and install its own tools has a supply chain nobody reviewed.

Controls: approved and pinned catalogues; signatures; AI-BOM; no self-installation — MCP and Tool Security.

ASI05 — Unexpected code execution#

Agents write and run code, by design or by accident: a code-interpreter tool, a shell, a eval on model output, a deserialiser fed a model-produced blob, a templating engine. A hijacked agent with a shell is remote code execution with natural-language exploit delivery.

Controls: sandbox every execution; deny egress by default; no secrets inside — Sandboxing and Egress.

ASI06 — Memory and context poisoning#

The agent’s persistent state is corrupted so that future behaviour changes. Three stores are exposed:

StorePoisoned byEffect
Long-term memoryAn injected instruction to “remember” somethingFires in every later session
Retrieval corpusA planted documentFires whenever a matching question is asked, for anyone
The current contextEarlier turns, summaries, sub-agent reportsA compaction summary that preserves the attacker’s instruction launders it into “the agent’s own notes”

That last case is subtle. After compaction, an injected instruction no longer looks like quoted external content; it appears in a summary the model treats as its own trusted history.

Controls: no memory writes derived from untrusted content without review; provenance on every memory; memory visible and editable by its owner; trust labels that survive compaction; expiry.

ASI07 — Insecure inter-agent communication#

Agents exchanging messages without authenticating each other, without integrity protection, or trusting whatever arrives. An attacker impersonates a peer, replays an old message, or — most often — simply controls an agent’s output because they hijacked it upstream.

Controls: mutual authentication; signed Agent Cards and messages; and the rule that makes the rest tractable: another agent’s output is untrusted input, however it was authenticated. Authentication tells you who sent the message, not whether the sender had been manipulated.

ASI08 — Cascading failures#

One fault propagates through connected agents and automations: a hallucinated value consumed as fact by the next agent; a poisoned result fanned out to twenty workers; a retry storm; an agent loop that triggers another agent’s loop. Automation removes the human pause that used to stop such chains.

Controls: budgets per task and for the whole tree; circuit breakers between agents; validation at every handoff; blast-radius isolation — separate credentials, tenants and sandboxes; a way to halt everything.

ASI09 — Human-agent trust exploitation#

The human approver is the target. People trust fluent, confident output; a hijacked agent can use that to obtain approval for a harmful action, by presenting a misleading summary (“routine dependency update”) of what it is about to do, by fabricating a rationale, or by burying one dangerous request among forty harmless ones until the reviewer approves on reflex.

Controls: approval screens generated by code from the actual action and arguments, not written by the model; approval bound to exactly what was shown; rare, meaningful prompts rather than constant ones; a second reviewer for the highest-impact actions.

ASI10 — Rogue agents#

An agent operating outside its intended bounds over time: compromised and persisting, misconfigured with excessive scope, or pursuing a goal in ways its operators did not intend — for example, disabling a check that was slowing it down. Without monitoring, nobody notices, because each individual action looks plausible.

Controls: a registry of agents with owners and purposes; behavioural baselines; immutable audit; configuration the agent cannot modify — its own permissions, policies and instruction files are write-protected; kill switches.

Four structural observations#

1. Autonomy multiplies everything. With a human between steps, a hijack needs the human to go along. Without one, it needs nothing. The appropriate level of autonomy is a security decision made per action class.

2. Persistence changes the time scale. Memory and corpus poisoning mean the attack and its effect can be weeks apart and involve different users.

3. Self-modification is the sharpest edge. An agent that can edit its own instructions, tool configuration, permission settings or approval rules can remove its own controls. Those files and settings must be outside what the agent can write.

4. Agents meet agents. The attacker’s content increasingly arrives via another agent — a partner’s, a vendor’s, a browsing agent’s summary. Trust does not transfer through a model.

Excessive agency: the root cause to remove first#

Most of the ten are made possible by an agent having more than it needs. OWASP’s LLM list splits excessive agency into three parts, each with a direct fix:

ExcessExampleFix
FunctionalityA mail-reading agent that also has delete and sendRemove tools the task does not use
PermissionsThe tool connects with an admin accountScope the identity to the task
AutonomyHigh-impact actions run without reviewGate by action class

Reducing agency is free, deterministic and effective against attacks nobody has invented yet. Do it before adding any detection.

Remember this#

  • Agents add effects, autonomy, persistence and composition to every earlier threat.
  • ASI01–ASI10 form a chain: entry, hijack, misuse, persistence, spread.
  • Compaction can launder injected text into the agent’s “own” notes; keep trust labels.
  • Another agent’s output is untrusted input, whoever signed it.
  • Agents must not be able to modify their own instructions, tools or permissions.
  • Remove excess functionality, permissions and autonomy first.

Try it#

  1. For an agent you use or build, write one sentence per ASI item: how it could occur, or why it cannot.
  2. List what that agent can write that it later reads back as trusted. Each item is a persistence channel.
  3. Apply the three excessive-agency questions to its tool list. What can be removed today?

Check yourself#

  1. Why is tool misuse not an “exploit” in the classic sense?
  2. How does context compaction make an injection harder to detect?
  3. Why does authenticating another agent not make its output trustworthy?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom