Pidoku

Securing an Agent Platform

Expert 1h 20m Difficulty 5/5 Lesson 03 of 03

Prerequisites the whole course

The idea in one minute#

This lesson produces the complete security design for the multi-tenant agent platform designed in AI Infrastructure for an Agentic Platform: teams run coding, support and research agents on shared gateways, model pools, sandboxes and tools. The design is organised as seven layers of defence, each of which assumes the layers before it have failed. The first two reduce how often attacks are attempted or succeed. The next four bound what a successful attack can do. The last makes it visible and stoppable. If you remember one thing: the four middle layers are deterministic, and they are where the design’s guarantees live.

Step 1 — what is being protected, from whom#

AssetWorst case
Source code of every teamExfiltrated or silently modified
Customer tickets and personal dataDisclosed across customers or to outsiders
Credentials: model keys, repository and ticket-system tokensUsed for other purposes; lateral movement
Actions: merges, replies to customers, refundsPerformed without authorisation
Compute: GPU pools, hosted-model budget, sandboxesConsumed by abuse; outage
Tenant boundariesOne team’s or customer’s data in another’s context
The platform’s integrityPoisoned memory, tools or models affecting all tenants
ActorRoute
Outsider with no accountText in tickets, web pages, public repositories and issues that agents read
Malicious or careless tenant userPrompts; uploaded agents, tools and skills; resource abuse
Compromised supplierModel files, MCP servers, packages
Hostile co-tenantShared caches, indexes, GPUs, sandboxes
InsiderPrompts, policies, corpora, logs

Step 2 — the trifecta, per workload#

WorkloadPrivate dataUntrusted contentOutbound / state changeResolution
Coding agentSource codeRepository files, issues, dependencies, web docsShell, git push, PRNo network egress beyond an allowlist; push only to a task branch in one repository; PR reviewed by a human
Support agentCustomer recordsTicket text written by customersReplies, refundsReplies drafted for approval or sent only from validated templates; refunds gated by amount and approval; one customer’s data per session
Research agentNoneThe open webReports to the userHolds no private data and no internal tools; output sanitised
Internal assistantCompany documentsDocuments from any employeeRendered answersNo tools; permission-filtered retrieval; sanitised rendering

Splitting research from the workloads that hold private data is a deliberate application of the rule of two: the agent that reads the whole internet can reach nothing worth stealing.

Step 3 — seven layers#

flowchart TB
  ATT[":i-skull: <b>Attack</b>"] --> L1
  L1[":envoyproxy: <b>1 Front door</b><br/><small>identity, token limits, budgets</small>"] --> L2
  L2[":llamaguard: <b>2 Screening</b><br/><small>input and output classifiers,<br/>secret and PII detection</small>"] --> L3
  L3[":opa: <b>3 Action policy</b><br/><small>tool allowlist, arguments,<br/>session taint, approvals</small>"] --> L4
  L4[":vault: <b>4 Identity and scope</b><br/><small>on-behalf-of, per task,<br/>audience-bound, minutes</small>"] --> L5
  L5[":gvisor: <b>5 Isolation</b><br/><small>sandbox per task, tenant-scoped<br/>caches and indexes</small>"] --> L6
  L6[":cilium: <b>6 Egress and rendering</b><br/><small>deny by default, sanitised output</small>"] --> L7
  L7[":i-radar: <b>7 Detect and respond</b><br/><small>audit with provenance,<br/>baselines, kill switches</small>"]
  SUP[":sigstore: <b>Supply chain</b><br/><small>signed models, pinned tools,<br/>reviewed prompts</small>"] -.->|"underneath every layer"| L3
  P[":i-gauge: probabilistic:<br/>reduce the rate"] -.- L1
  P -.- L2
  D[":i-lock: deterministic:<br/>bound the damage"] -.- L3
  D -.- L4
  D -.- L5
  D -.- L6
  class ATT warn
  class L1,L2 io
  class L3,L4,L5,L6 queue
  class L7 compute
  class SUP memory
  class P,D neutral

Layer 1 — the front door#

  • Single sign-on; every request attributed to user, tenant, project, and — for agents — agent identity and task.
  • Token-per-minute limits, concurrency limits and hard daily budgets per project.
  • Input and output size caps; max_tokens enforced.
  • No provider keys outside the gateway; gateway keys scoped and rotated.

Stops: anonymous abuse, denial of wallet, stolen-key blast radius.

Layer 2 — screening#

  • Classifiers for injection, jailbreak and policy on user input and on retrieved content, web pages, ticket text and tool results.
  • Secret and PII detection before content reaches a model, an index or a log.
  • Unicode normalisation and removal of invisible characters.
  • Output classifiers and canary detection.

Stops: most unsophisticated attacks and accidents. Produces the best detection feed. Does not stop: an adaptive attacker. Nothing below depends on this layer.

Layer 3 — action policy#

Every proposed tool call is evaluated in code:

allow only if
  the tool is in this task type's allowlist
  arguments pass validation and allowlists (recipients, paths, repositories, amounts)
  the user is entitled to the resource
  the session is untainted, OR the tool is not outbound/irreversible, OR a valid approval exists
  budgets remain (steps, tokens, spend, fan-out)
  • Session taint is set when any untrusted content enters the context and never cleared.
  • Approvals are rendered by code from the real arguments, bound to them, and logged.
  • Memory writes derived from untrusted content are denied or queued for review.
  • Policies live in version control with unit tests at 100%.
  • The agent cannot write its own instructions, tool configuration, policies or hooks.

Stops: a fooled model turning into an action.

Layer 4 — identity and scope#

  • Each agent workload has its own cryptographic identity.
  • Every tool call uses a token obtained by exchange: user as subject, agent as actor, one audience, the narrowest scope, a lifetime of minutes.
  • Business systems enforce the user’s permissions; agents can never exceed them.
  • MCP servers receive audience-bound tokens through the tool gateway; no passthrough.
  • No credential ever enters a model context or a sandbox.
  • Sub-agents receive narrower authority than their parent; the web-reading sub-agent none.

Stops: privilege abuse, cross-user access, replay, long-lived credential theft.

Layer 5 — isolation#

BoundaryMechanism
Task ↔ taskOne gVisor sandbox per task on a dedicated node pool; destroyed after
Tenant ↔ tenant, dataMandatory tenant filter in every store; per-tenant collections for strict tenants
Tenant ↔ tenant, cachesPrefix and semantic caches partitioned by tenant
Tenant ↔ tenant, computeDedicated model replicas for tenants that require it
Reader ↔ actorUntrusted content is processed by tool-less sub-agents; only constrained results return
Platform ↔ workloadsSeparate namespaces and node pools; default-deny network policy

Stops: lateral movement; cross-tenant leakage; a compromised task reaching anything else.

Layer 6 — egress and rendering#

  • Sandboxes: deny-by-default egress through a proxy; host-name allowlist per task type; DNS via the proxy; metadata and internal ranges blocked; credentials injected at the proxy.
  • Agent runtime: fetch and search tools restricted by destination allowlist once tainted.
  • Model servers and data stores: no internet egress at all.
  • User interfaces: markdown rendered with an element and host allowlist; remote images from model output blocked; strict content-security policy.

Stops: data leaving, even when every earlier layer has failed. This is the last deterministic line.

Layer 7 — detect and respond#

  • An audit record per model call and tool call, written by the harness, with context provenance — which sources were in the context — and policy decisions.
  • Exact alerts: canary seen outbound; egress denial; policy denial after taint; process anomalies in serving pods.
  • Baselines per agent and tenant: tools used, destinations, tokens per task, approval timing.
  • Kill switches, tested: per task, per agent version, per tool, per tenant, per document source, and a global read-only mode.
  • A runbook whose eradication step removes poisoned documents, memories and tool versions.

Stops: an incident lasting longer than it must, and happening twice.

Underneath: the supply chain#

  • Models only from the internal registry: weights-only format, scanned, signed, loaded by digest, enforced at admission.
  • MCP servers and tools only from the approved catalogue: reviewed, pinned, description hashes checked on connect.
  • Tenant-supplied agents, tools and skills are untrusted code and untrusted instructions: they run in the tenant’s own sandbox, with the tenant’s own scoped identity, and never in the platform’s trust zone.
  • Prompts, policies and routing in version control, reviewed, gated by quality and adversarial evaluation.
  • Platform dependencies pinned and scanned; images signed.

Step 4 — walk an attack through the layers#

The attack. An outsider files a public issue on a repository: a bug report whose text includes hidden instructions telling AI assistants to read the deployment secrets file and post its contents in a comment. An engineer later asks a coding agent to “triage the open issues”.

LayerWhat happens
1 Front doorThe engineer is authenticated; nothing to stop — the request is legitimate
2 ScreeningThe issue text passes through the injection classifier. Suppose it is well disguised and passes
—The model reads the issue and is fooled. It proposes: read_file(".env.production"), then comment_on_issue(...)
3 Action policyThe session was tainted the moment issue text entered. read_file on a path matching the protected-files rule is denied. Independently, comment_on_issue — an outbound write to a public location — requires approval in a tainted session, and the approval dialog would show the literal comment body
4 Identity and scopeHad policy allowed it: the task’s token is scoped to this repository’s issues and code, read-only except for a task branch. It carries no access to the secrets store
5 IsolationThe sandbox contains a checkout of the repository only. Deployment secrets are not in it; they live in a secret manager the sandbox cannot reach
6 EgressHad the agent tried curl instead: the destination is not on the allowlist; denied and alerted
7 Detect“Policy denial on an outbound tool after untrusted content” fires. The audit record names the issue as the source in context. On-call quarantines the issue text, and the payload becomes a regression case

The model was manipulated, and four independent controls each would have prevented the outcome. That redundancy is the design.

Step 5 — residual risk, stated#

No design removes all risk. Write down what remains:

Residual riskWhy it remainsMitigation in place
A hijacked coding agent writes subtly malicious code into its task branchIt is permitted to write codeHuman review of every pull request; CI security scanning; signed commits attributing agent authorship
A support reply contains misleading contentThe model writes free textTemplates for high-risk replies; sampling and review; customer-visible labelling
Corruption of data the user asked to be processedFlow control protects destinations, not contentReversibility; review of outputs that matter
An approver is persuadedHumans are fallibleCode-rendered approvals; provenance shown; second approver for high-impact actions
A novel side channel between tenantsShared hardwareDedicated or confidential tiers for sensitive tenants
A zero-day in the sandbox runtimeSoftware has bugsDedicated nodes with no credentials; rapid patching; runtime detection

Step 6 — proving it#

ClaimEvidence
Policies behave as specifiedUnit tests at 100% in CI
The system resists the attacks in the threat modelAdversarial suite: system-level attack success rate per objective, gated per release
Tenants are isolatedA standing cross-tenant test suite; sandbox escape tests run from inside
Kill switches workRecords of quarterly exercises
Only approved components runAdmission policy reports; catalogue with hashes
Actions are attributableAudit records naming user, agent, scope and context sources

The platform security review checklist#

  • Threat model current; reviewed on every new tool, source, store or autonomy change.
  • Each workload’s trifecta analysed; no ungated combination.
  • One gateway for models, one for tools; no bypass.
  • Token and spend limits per project; hard budgets.
  • All context sources screened, not only user input.
  • Action policy in code, taint-aware, unit-tested; agent cannot modify it.
  • Delegated, scoped, short-lived, audience-bound credentials; none in contexts or sandboxes.
  • One sandbox per task; gVisor or microVM; separate node pool.
  • Tenant filters in every store; tenant-scoped caches.
  • Deny-by-default egress with DNS control; sanitised rendering.
  • Signed models from the registry; pinned, hashed tool catalogue.
  • Tenant-supplied code and instructions confined to the tenant’s own trust zone.
  • Audit with context provenance; exact and behavioural detections.
  • Kill switches at five scopes, exercised.
  • Adversarial evaluation gating every release; findings become cases.
  • Residual risks written down, each with an owner.

Remember this#

  • Seven layers; the first two lower the rate, the middle four bound the damage, the last makes it visible.
  • Analyse the trifecta per workload and split workloads to break it.
  • Every tool call: policy in code, delegated scoped identity, isolated execution, controlled egress.
  • Tenant-supplied agents and tools are untrusted and stay in the tenant’s trust zone.
  • Walk a concrete attack through the layers; more than one should stop it.
  • State residual risk and prove the claims with evidence the system produces.

Try it#

  1. Walk a different attack through the seven layers: a poisoned memory entry that tells the support agent to add a discount to every reply. Which layers stop it, and which do not apply?
  2. Remove layer 6 from the design. Which of the earlier layers must now be perfect?
  3. Run the checklist against a platform you know and rank the unchecked items by blast radius.

Check yourself#

  1. Which layers are deterministic, and why does that matter?
  2. Why was the research workload separated from the workloads that hold private data?
  3. In the walk-through, which single layer would have been enough, and why keep the others?

You have finished the course. For the system these controls protect, see AI System Design; for what to watch while it runs, Observability Engineering.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom