Pidoku

Monitoring and Response

Intermediate 45 min Difficulty 3/5 Lesson 07 of 07

Prerequisites Runtime Guardrails

The idea in one minute#

Prevention will sometimes fail, so you need to see an attack on an AI system and stop it. Seeing requires an audit trail that records not only what was done but what the model had read when it decided to do it — without that, you cannot tell a hijacked agent from a buggy one. Stopping requires controls that exist before the incident: a way to halt a task, an agent, a tool or a tenant in seconds, and to revoke what it was holding. AI incidents also have a clean-up step that classic ones lack: finding and removing the poisoned document, memory or tool that caused it, or the attack simply runs again.

A picture#

flowchart LR
  subgraph SIGNALS["Signals"]
    direction LR
    S1[":envoyproxy: <b>Gateway</b><br/><small>usage, limits, classifier hits</small>"]
    S2[":i-bot: <b>Agent runtime</b><br/><small>tool calls, policy denials,<br/>context provenance</small>"]
    S3[":cilium: <b>Network</b><br/><small>egress denials</small>"]
    S4[":falco: <b>Runtime</b><br/><small>unexpected processes</small>"]
  end
  SIGNALS --> PIPE[":opentelemetry: <b>Telemetry pipeline</b>"]
  PIPE --> STORE[(":clickhouse: <b>Audit store</b><br/><small>immutable, joined by trace ID</small>")]
  STORE --> DET[":i-radar: <b>Detections</b><br/><small>rules + behaviour baselines</small>"]
  DET --> ONC[":pagerduty: <b>On-call</b>"]
  ONC --> KILL[":i-ban: <b>Contain</b><br/><small>stop task, agent, tool, tenant;<br/>revoke tokens</small>"]
  KILL --> CLEAN[":i-recycle: <b>Eradicate</b><br/><small>remove poisoned source,<br/>memory, tool version</small>"]
  CLEAN --> LEARN[":i-list-checks: <b>Learn</b><br/><small>new test case,<br/>threat-model update</small>"]
  class S1,S2,S3,S4 io
  class PIPE queue
  class STORE memory
  class DET compute
  class ONC neutral
  class KILL,CLEAN warn
  class LEARN neutral

How it really works#

The audit record#

One record per model call and per tool call, joined by a trace ID across the whole task:

identity      tenant, user, agent name and version, task ID, session ID
model call    model and version, prompt template version, token counts, finish reason
context       for each item in the context: source type, source ID, trust label, hash
tool call     tool, server, arguments (redacted where needed), result status and size
decision      policy rule evaluated, verdict, approval (who, when, what exactly)
credentials   scope and audience of the token used; not the token
guardrails    classifier names, scores, verdicts
effects       resources read and written; outbound destinations

The context block is what distinguishes an AI audit trail. With it, an investigator can answer the first question of every agent incident: what text caused this? Store hashes and source references for everything, and full content only where retention policy allows — prompts and completions are sensitive data, kept separately with restricted access.

Records must be tamper-resistant and written by the harness, not the model. A model asked to “log what you did” produces a story; the harness produces a fact.

What to detect#

SignalWhy it matters
A policy denial on an outbound tool shortly after untrusted content entered the contextThe signature of an injection attempt that the action rail stopped
First-time use: a tool, destination domain or resource an agent has not used beforeHijacked agents do new things
Egress denials from sandboxes or agent podsSomething tried to reach an address outside the allowlist
Canary hits: a planted token appearing in output, a URL or an outbound requestDirect evidence of exfiltration
Classifier hits clustering on one user, document or sourceAn attacker iterating, or a poisoned document being retrieved repeatedly
Memory writes triggered by content rather than by the userPossible persistence
Approval anomalies: bursts of requests; approvals faster than a human can readConsent fatigue being exploited, or automation of the approver
Budget hits and loopsRunaway or resource abuse
Retrieval anomalies: a new document suddenly ranking first for sensitive queriesCorpus poisoning
Spend spikes per key, tenant or agentStolen credentials or denial of wallet
Hidden-context probes: many variations asking for the system promptReconnaissance
Process or network surprises in serving podsInfrastructure compromise

Two kinds of rule are needed. Exact rules fire on events that should never happen: a canary leaving, an egress denial, a model server spawning a shell. Behavioural baselines compare an agent or tenant with its own history — tools used, destinations contacted, tokens per task — and flag departures. Agents are more regular than people, which makes baselining them practical.

MITRE ATLAS techniques make a good coverage checklist: for each technique relevant to your system, name the log line that would reveal it.

Containment: built before it is needed#

ScopeControl
One taskCancel; destroy its sandbox; revoke its tokens
One agent or agent versionDisable at the runtime; roll back the prompt, tool set or model version
One tool or MCP serverRemove from the catalogue at the tool gateway; all agents lose it at once
One tenant or userSuspend keys at the gateway
One document or sourceQuarantine it in the index; exclude the source from retrieval
EverythingA global switch that moves agents to read-only or suggest-only mode

Each of these should be a one-step operation that on-call can perform at 03:00 without writing code, and each should be exercised — a kill switch never tested is a hope. Because per-task credentials are short-lived and scoped, containment mostly means “stop issuing new ones”, which is far faster than rotating a shared secret.

The response, step by step#

1. Triage. Was an action taken, or only attempted? Was data sent out? Which identity’s authority was used? The audit trail answers these in minutes if it exists.

2. Contain. Narrowest scope that stops the harm: the task, then the tool, then the agent.

3. Find the cause. Follow the trace back from the malicious action to the context that preceded it, and identify the source: a document, an email, a web page, a tool result, a memory entry, a tool description.

4. Eradicate. This is the AI-specific step. Remove the cause everywhere:

  • delete or quarantine the poisoned document and its chunks;
  • purge affected memory entries — and check what else that session wrote;
  • pin or remove the tool or MCP server version;
  • search for the same payload elsewhere in the corpus, inboxes and repositories, since attackers plant many copies;
  • check whether the payload spread: did the agent write it into files, messages or other agents’ inputs?

5. Assess the damage. From the effects recorded: what was read, written and sent, with whose authority. Notify as contracts and law require.

6. Recover. Restore changed data; re-enable capabilities in stages.

7. Learn. The payload becomes a permanent case in the adversarial suite. The threat model is updated. If a probabilistic control was the only thing in the way, add a deterministic one.

What makes AI incidents awkward#

  • Non-reproducibility. The same input may not misbehave on replay. Investigate from the recorded trace, not from re-running it.
  • Ambiguous intent. A wrong action can be an attack, a model error or a bad prompt. The context provenance usually settles it: was there instruction-like text from an untrusted source?
  • Delayed triggers. Poison planted weeks ago in memory or a document fires today. Keep provenance long enough to trace it.
  • Reporting channels. Users and researchers who find a prompt-injection path need somewhere to report it; add AI issues to your vulnerability-disclosure policy.

Protecting the monitoring itself#

Telemetry about prompts is a concentration of sensitive data and an attack surface:

  • Content logging is off by default, opt-in, redacted, short-lived and access-audited.
  • Log viewers and dashboards that render model output need the same output sanitising as the product — an injected payload that executes in an analyst’s browser is an old trick with a new source.
  • Analysts’ AI assistants that read logs are agents reading untrusted content. Apply this whole course to them.

Remember this#

  • Record what the model had read, not only what it did. Provenance is the key to every investigation.
  • Detect with exact rules for never-events and baselines for behaviour.
  • Containment controls — per task, agent, tool, tenant, source — exist and are tested in advance.
  • Eradication means removing the poisoned source everywhere, including memory and copies.
  • Every incident ends as a test case and a threat-model change.

Try it#

  1. Write the audit record for one tool call in a system you know. Which fields are missing?
  2. Choose three signals from the table and write the exact query or rule for each.
  3. Run a tabletop: an agent emailed a customer list to an outside address. Walk the seven steps and note where you would be stuck today.

Check yourself#

  1. Why must the audit trail be written by the harness rather than described by the model?
  2. What is the AI-specific step in incident response, and why is it necessary?
  3. Why are short-lived per-task credentials an advantage during containment?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom