The idea in one minute#
Prevention will sometimes fail, so you need to see an attack on an AI system and stop it. Seeing requires an audit trail that records not only what was done but what the model had read when it decided to do it — without that, you cannot tell a hijacked agent from a buggy one. Stopping requires controls that exist before the incident: a way to halt a task, an agent, a tool or a tenant in seconds, and to revoke what it was holding. AI incidents also have a clean-up step that classic ones lack: finding and removing the poisoned document, memory or tool that caused it, or the attack simply runs again.
A picture#
flowchart LR
subgraph SIGNALS["Signals"]
direction LR
S1[":envoyproxy: <b>Gateway</b><br/><small>usage, limits, classifier hits</small>"]
S2[":i-bot: <b>Agent runtime</b><br/><small>tool calls, policy denials,<br/>context provenance</small>"]
S3[":cilium: <b>Network</b><br/><small>egress denials</small>"]
S4[":falco: <b>Runtime</b><br/><small>unexpected processes</small>"]
end
SIGNALS --> PIPE[":opentelemetry: <b>Telemetry pipeline</b>"]
PIPE --> STORE[(":clickhouse: <b>Audit store</b><br/><small>immutable, joined by trace ID</small>")]
STORE --> DET[":i-radar: <b>Detections</b><br/><small>rules + behaviour baselines</small>"]
DET --> ONC[":pagerduty: <b>On-call</b>"]
ONC --> KILL[":i-ban: <b>Contain</b><br/><small>stop task, agent, tool, tenant;<br/>revoke tokens</small>"]
KILL --> CLEAN[":i-recycle: <b>Eradicate</b><br/><small>remove poisoned source,<br/>memory, tool version</small>"]
CLEAN --> LEARN[":i-list-checks: <b>Learn</b><br/><small>new test case,<br/>threat-model update</small>"]
class S1,S2,S3,S4 io
class PIPE queue
class STORE memory
class DET compute
class ONC neutral
class KILL,CLEAN warn
class LEARN neutralHow it really works#
The audit record#
One record per model call and per tool call, joined by a trace ID across the whole task:
identity tenant, user, agent name and version, task ID, session ID
model call model and version, prompt template version, token counts, finish reason
context for each item in the context: source type, source ID, trust label, hash
tool call tool, server, arguments (redacted where needed), result status and size
decision policy rule evaluated, verdict, approval (who, when, what exactly)
credentials scope and audience of the token used; not the token
guardrails classifier names, scores, verdicts
effects resources read and written; outbound destinationsThe context block is what distinguishes an AI audit trail. With it, an investigator can answer the first question of every agent incident: what text caused this? Store hashes and source references for everything, and full content only where retention policy allows — prompts and completions are sensitive data, kept separately with restricted access.
Records must be tamper-resistant and written by the harness, not the model. A model asked to “log what you did” produces a story; the harness produces a fact.
What to detect#
| Signal | Why it matters |
|---|---|
| A policy denial on an outbound tool shortly after untrusted content entered the context | The signature of an injection attempt that the action rail stopped |
| First-time use: a tool, destination domain or resource an agent has not used before | Hijacked agents do new things |
| Egress denials from sandboxes or agent pods | Something tried to reach an address outside the allowlist |
| Canary hits: a planted token appearing in output, a URL or an outbound request | Direct evidence of exfiltration |
| Classifier hits clustering on one user, document or source | An attacker iterating, or a poisoned document being retrieved repeatedly |
| Memory writes triggered by content rather than by the user | Possible persistence |
| Approval anomalies: bursts of requests; approvals faster than a human can read | Consent fatigue being exploited, or automation of the approver |
| Budget hits and loops | Runaway or resource abuse |
| Retrieval anomalies: a new document suddenly ranking first for sensitive queries | Corpus poisoning |
| Spend spikes per key, tenant or agent | Stolen credentials or denial of wallet |
| Hidden-context probes: many variations asking for the system prompt | Reconnaissance |
| Process or network surprises in serving pods | Infrastructure compromise |
Two kinds of rule are needed. Exact rules fire on events that should never happen: a canary leaving, an egress denial, a model server spawning a shell. Behavioural baselines compare an agent or tenant with its own history — tools used, destinations contacted, tokens per task — and flag departures. Agents are more regular than people, which makes baselining them practical.
MITRE ATLAS techniques make a good coverage checklist: for each technique relevant to your system, name the log line that would reveal it.
Containment: built before it is needed#
| Scope | Control |
|---|---|
| One task | Cancel; destroy its sandbox; revoke its tokens |
| One agent or agent version | Disable at the runtime; roll back the prompt, tool set or model version |
| One tool or MCP server | Remove from the catalogue at the tool gateway; all agents lose it at once |
| One tenant or user | Suspend keys at the gateway |
| One document or source | Quarantine it in the index; exclude the source from retrieval |
| Everything | A global switch that moves agents to read-only or suggest-only mode |
Each of these should be a one-step operation that on-call can perform at 03:00 without writing code, and each should be exercised — a kill switch never tested is a hope. Because per-task credentials are short-lived and scoped, containment mostly means “stop issuing new ones”, which is far faster than rotating a shared secret.
The response, step by step#
1. Triage. Was an action taken, or only attempted? Was data sent out? Which identity’s authority was used? The audit trail answers these in minutes if it exists.
2. Contain. Narrowest scope that stops the harm: the task, then the tool, then the agent.
3. Find the cause. Follow the trace back from the malicious action to the context that preceded it, and identify the source: a document, an email, a web page, a tool result, a memory entry, a tool description.
4. Eradicate. This is the AI-specific step. Remove the cause everywhere:
- delete or quarantine the poisoned document and its chunks;
- purge affected memory entries — and check what else that session wrote;
- pin or remove the tool or MCP server version;
- search for the same payload elsewhere in the corpus, inboxes and repositories, since attackers plant many copies;
- check whether the payload spread: did the agent write it into files, messages or other agents’ inputs?
5. Assess the damage. From the effects recorded: what was read, written and sent, with whose authority. Notify as contracts and law require.
6. Recover. Restore changed data; re-enable capabilities in stages.
7. Learn. The payload becomes a permanent case in the adversarial suite. The threat model is updated. If a probabilistic control was the only thing in the way, add a deterministic one.
What makes AI incidents awkward#
- Non-reproducibility. The same input may not misbehave on replay. Investigate from the recorded trace, not from re-running it.
- Ambiguous intent. A wrong action can be an attack, a model error or a bad prompt. The context provenance usually settles it: was there instruction-like text from an untrusted source?
- Delayed triggers. Poison planted weeks ago in memory or a document fires today. Keep provenance long enough to trace it.
- Reporting channels. Users and researchers who find a prompt-injection path need somewhere to report it; add AI issues to your vulnerability-disclosure policy.
Protecting the monitoring itself#
Telemetry about prompts is a concentration of sensitive data and an attack surface:
- Content logging is off by default, opt-in, redacted, short-lived and access-audited.
- Log viewers and dashboards that render model output need the same output sanitising as the product — an injected payload that executes in an analyst’s browser is an old trick with a new source.
- Analysts’ AI assistants that read logs are agents reading untrusted content. Apply this whole course to them.
Remember this#
- Record what the model had read, not only what it did. Provenance is the key to every investigation.
- Detect with exact rules for never-events and baselines for behaviour.
- Containment controls — per task, agent, tool, tenant, source — exist and are tested in advance.
- Eradication means removing the poisoned source everywhere, including memory and copies.
- Every incident ends as a test case and a threat-model change.
Try it#
- Write the audit record for one tool call in a system you know. Which fields are missing?
- Choose three signals from the table and write the exact query or rule for each.
- Run a tabletop: an agent emailed a customer list to an outside address. Walk the seven steps and note where you would be stuck today.
Check yourself#
- Why must the audit trail be written by the harness rather than described by the model?
- What is the AI-specific step in incident response, and why is it necessary?
- Why are short-lived per-task credentials an advantage during containment?