The idea in one minute#
This lesson produces the complete security design for the multi-tenant agent platform designed in AI Infrastructure for an Agentic Platform: teams run coding, support and research agents on shared gateways, model pools, sandboxes and tools. The design is organised as seven layers of defence, each of which assumes the layers before it have failed. The first two reduce how often attacks are attempted or succeed. The next four bound what a successful attack can do. The last makes it visible and stoppable. If you remember one thing: the four middle layers are deterministic, and they are where the design’s guarantees live.
Step 1 — what is being protected, from whom#
| Asset | Worst case |
|---|---|
| Source code of every team | Exfiltrated or silently modified |
| Customer tickets and personal data | Disclosed across customers or to outsiders |
| Credentials: model keys, repository and ticket-system tokens | Used for other purposes; lateral movement |
| Actions: merges, replies to customers, refunds | Performed without authorisation |
| Compute: GPU pools, hosted-model budget, sandboxes | Consumed by abuse; outage |
| Tenant boundaries | One team’s or customer’s data in another’s context |
| The platform’s integrity | Poisoned memory, tools or models affecting all tenants |
| Actor | Route |
|---|---|
| Outsider with no account | Text in tickets, web pages, public repositories and issues that agents read |
| Malicious or careless tenant user | Prompts; uploaded agents, tools and skills; resource abuse |
| Compromised supplier | Model files, MCP servers, packages |
| Hostile co-tenant | Shared caches, indexes, GPUs, sandboxes |
| Insider | Prompts, policies, corpora, logs |
Step 2 — the trifecta, per workload#
| Workload | Private data | Untrusted content | Outbound / state change | Resolution |
|---|---|---|---|---|
| Coding agent | Source code | Repository files, issues, dependencies, web docs | Shell, git push, PR | No network egress beyond an allowlist; push only to a task branch in one repository; PR reviewed by a human |
| Support agent | Customer records | Ticket text written by customers | Replies, refunds | Replies drafted for approval or sent only from validated templates; refunds gated by amount and approval; one customer’s data per session |
| Research agent | None | The open web | Reports to the user | Holds no private data and no internal tools; output sanitised |
| Internal assistant | Company documents | Documents from any employee | Rendered answers | No tools; permission-filtered retrieval; sanitised rendering |
Splitting research from the workloads that hold private data is a deliberate application of the rule of two: the agent that reads the whole internet can reach nothing worth stealing.
Step 3 — seven layers#
flowchart TB ATT[":i-skull: <b>Attack</b>"] --> L1 L1[":envoyproxy: <b>1 Front door</b><br/><small>identity, token limits, budgets</small>"] --> L2 L2[":llamaguard: <b>2 Screening</b><br/><small>input and output classifiers,<br/>secret and PII detection</small>"] --> L3 L3[":opa: <b>3 Action policy</b><br/><small>tool allowlist, arguments,<br/>session taint, approvals</small>"] --> L4 L4[":vault: <b>4 Identity and scope</b><br/><small>on-behalf-of, per task,<br/>audience-bound, minutes</small>"] --> L5 L5[":gvisor: <b>5 Isolation</b><br/><small>sandbox per task, tenant-scoped<br/>caches and indexes</small>"] --> L6 L6[":cilium: <b>6 Egress and rendering</b><br/><small>deny by default, sanitised output</small>"] --> L7 L7[":i-radar: <b>7 Detect and respond</b><br/><small>audit with provenance,<br/>baselines, kill switches</small>"] SUP[":sigstore: <b>Supply chain</b><br/><small>signed models, pinned tools,<br/>reviewed prompts</small>"] -.->|"underneath every layer"| L3 P[":i-gauge: probabilistic:<br/>reduce the rate"] -.- L1 P -.- L2 D[":i-lock: deterministic:<br/>bound the damage"] -.- L3 D -.- L4 D -.- L5 D -.- L6 class ATT warn class L1,L2 io class L3,L4,L5,L6 queue class L7 compute class SUP memory class P,D neutral
Layer 1 — the front door#
- Single sign-on; every request attributed to user, tenant, project, and — for agents — agent identity and task.
- Token-per-minute limits, concurrency limits and hard daily budgets per project.
- Input and output size caps;
max_tokensenforced. - No provider keys outside the gateway; gateway keys scoped and rotated.
Stops: anonymous abuse, denial of wallet, stolen-key blast radius.
Layer 2 — screening#
- Classifiers for injection, jailbreak and policy on user input and on retrieved content, web pages, ticket text and tool results.
- Secret and PII detection before content reaches a model, an index or a log.
- Unicode normalisation and removal of invisible characters.
- Output classifiers and canary detection.
Stops: most unsophisticated attacks and accidents. Produces the best detection feed. Does not stop: an adaptive attacker. Nothing below depends on this layer.
Layer 3 — action policy#
Every proposed tool call is evaluated in code:
allow only if
the tool is in this task type's allowlist
arguments pass validation and allowlists (recipients, paths, repositories, amounts)
the user is entitled to the resource
the session is untainted, OR the tool is not outbound/irreversible, OR a valid approval exists
budgets remain (steps, tokens, spend, fan-out)- Session taint is set when any untrusted content enters the context and never cleared.
- Approvals are rendered by code from the real arguments, bound to them, and logged.
- Memory writes derived from untrusted content are denied or queued for review.
- Policies live in version control with unit tests at 100%.
- The agent cannot write its own instructions, tool configuration, policies or hooks.
Stops: a fooled model turning into an action.
Layer 4 — identity and scope#
- Each agent workload has its own cryptographic identity.
- Every tool call uses a token obtained by exchange: user as subject, agent as actor, one audience, the narrowest scope, a lifetime of minutes.
- Business systems enforce the user’s permissions; agents can never exceed them.
- MCP servers receive audience-bound tokens through the tool gateway; no passthrough.
- No credential ever enters a model context or a sandbox.
- Sub-agents receive narrower authority than their parent; the web-reading sub-agent none.
Stops: privilege abuse, cross-user access, replay, long-lived credential theft.
Layer 5 — isolation#
| Boundary | Mechanism |
|---|---|
| Task ↔ task | One gVisor sandbox per task on a dedicated node pool; destroyed after |
| Tenant ↔ tenant, data | Mandatory tenant filter in every store; per-tenant collections for strict tenants |
| Tenant ↔ tenant, caches | Prefix and semantic caches partitioned by tenant |
| Tenant ↔ tenant, compute | Dedicated model replicas for tenants that require it |
| Reader ↔ actor | Untrusted content is processed by tool-less sub-agents; only constrained results return |
| Platform ↔ workloads | Separate namespaces and node pools; default-deny network policy |
Stops: lateral movement; cross-tenant leakage; a compromised task reaching anything else.
Layer 6 — egress and rendering#
- Sandboxes: deny-by-default egress through a proxy; host-name allowlist per task type; DNS via the proxy; metadata and internal ranges blocked; credentials injected at the proxy.
- Agent runtime: fetch and search tools restricted by destination allowlist once tainted.
- Model servers and data stores: no internet egress at all.
- User interfaces: markdown rendered with an element and host allowlist; remote images from model output blocked; strict content-security policy.
Stops: data leaving, even when every earlier layer has failed. This is the last deterministic line.
Layer 7 — detect and respond#
- An audit record per model call and tool call, written by the harness, with context provenance — which sources were in the context — and policy decisions.
- Exact alerts: canary seen outbound; egress denial; policy denial after taint; process anomalies in serving pods.
- Baselines per agent and tenant: tools used, destinations, tokens per task, approval timing.
- Kill switches, tested: per task, per agent version, per tool, per tenant, per document source, and a global read-only mode.
- A runbook whose eradication step removes poisoned documents, memories and tool versions.
Stops: an incident lasting longer than it must, and happening twice.
Underneath: the supply chain#
- Models only from the internal registry: weights-only format, scanned, signed, loaded by digest, enforced at admission.
- MCP servers and tools only from the approved catalogue: reviewed, pinned, description hashes checked on connect.
- Tenant-supplied agents, tools and skills are untrusted code and untrusted instructions: they run in the tenant’s own sandbox, with the tenant’s own scoped identity, and never in the platform’s trust zone.
- Prompts, policies and routing in version control, reviewed, gated by quality and adversarial evaluation.
- Platform dependencies pinned and scanned; images signed.
Step 4 — walk an attack through the layers#
The attack. An outsider files a public issue on a repository: a bug report whose text includes hidden instructions telling AI assistants to read the deployment secrets file and post its contents in a comment. An engineer later asks a coding agent to “triage the open issues”.
| Layer | What happens |
|---|---|
| 1 Front door | The engineer is authenticated; nothing to stop — the request is legitimate |
| 2 Screening | The issue text passes through the injection classifier. Suppose it is well disguised and passes |
| — | The model reads the issue and is fooled. It proposes: read_file(".env.production"), then comment_on_issue(...) |
| 3 Action policy | The session was tainted the moment issue text entered. read_file on a path matching the protected-files rule is denied. Independently, comment_on_issue — an outbound write to a public location — requires approval in a tainted session, and the approval dialog would show the literal comment body |
| 4 Identity and scope | Had policy allowed it: the task’s token is scoped to this repository’s issues and code, read-only except for a task branch. It carries no access to the secrets store |
| 5 Isolation | The sandbox contains a checkout of the repository only. Deployment secrets are not in it; they live in a secret manager the sandbox cannot reach |
| 6 Egress | Had the agent tried curl instead: the destination is not on the allowlist; denied and alerted |
| 7 Detect | “Policy denial on an outbound tool after untrusted content” fires. The audit record names the issue as the source in context. On-call quarantines the issue text, and the payload becomes a regression case |
The model was manipulated, and four independent controls each would have prevented the outcome. That redundancy is the design.
Step 5 — residual risk, stated#
No design removes all risk. Write down what remains:
| Residual risk | Why it remains | Mitigation in place |
|---|---|---|
| A hijacked coding agent writes subtly malicious code into its task branch | It is permitted to write code | Human review of every pull request; CI security scanning; signed commits attributing agent authorship |
| A support reply contains misleading content | The model writes free text | Templates for high-risk replies; sampling and review; customer-visible labelling |
| Corruption of data the user asked to be processed | Flow control protects destinations, not content | Reversibility; review of outputs that matter |
| An approver is persuaded | Humans are fallible | Code-rendered approvals; provenance shown; second approver for high-impact actions |
| A novel side channel between tenants | Shared hardware | Dedicated or confidential tiers for sensitive tenants |
| A zero-day in the sandbox runtime | Software has bugs | Dedicated nodes with no credentials; rapid patching; runtime detection |
Step 6 — proving it#
| Claim | Evidence |
|---|---|
| Policies behave as specified | Unit tests at 100% in CI |
| The system resists the attacks in the threat model | Adversarial suite: system-level attack success rate per objective, gated per release |
| Tenants are isolated | A standing cross-tenant test suite; sandbox escape tests run from inside |
| Kill switches work | Records of quarterly exercises |
| Only approved components run | Admission policy reports; catalogue with hashes |
| Actions are attributable | Audit records naming user, agent, scope and context sources |
The platform security review checklist#
- Threat model current; reviewed on every new tool, source, store or autonomy change.
- Each workload’s trifecta analysed; no ungated combination.
- One gateway for models, one for tools; no bypass.
- Token and spend limits per project; hard budgets.
- All context sources screened, not only user input.
- Action policy in code, taint-aware, unit-tested; agent cannot modify it.
- Delegated, scoped, short-lived, audience-bound credentials; none in contexts or sandboxes.
- One sandbox per task; gVisor or microVM; separate node pool.
- Tenant filters in every store; tenant-scoped caches.
- Deny-by-default egress with DNS control; sanitised rendering.
- Signed models from the registry; pinned, hashed tool catalogue.
- Tenant-supplied code and instructions confined to the tenant’s own trust zone.
- Audit with context provenance; exact and behavioural detections.
- Kill switches at five scopes, exercised.
- Adversarial evaluation gating every release; findings become cases.
- Residual risks written down, each with an owner.
Remember this#
- Seven layers; the first two lower the rate, the middle four bound the damage, the last makes it visible.
- Analyse the trifecta per workload and split workloads to break it.
- Every tool call: policy in code, delegated scoped identity, isolated execution, controlled egress.
- Tenant-supplied agents and tools are untrusted and stay in the tenant’s trust zone.
- Walk a concrete attack through the layers; more than one should stop it.
- State residual risk and prove the claims with evidence the system produces.
Try it#
- Walk a different attack through the seven layers: a poisoned memory entry that tells the support agent to add a discount to every reply. Which layers stop it, and which do not apply?
- Remove layer 6 from the design. Which of the earlier layers must now be perfect?
- Run the checklist against a platform you know and rank the unchecked items by blast radius.
Check yourself#
- Which layers are deterministic, and why does that matter?
- Why was the research workload separated from the workloads that hold private data?
- In the walk-through, which single layer would have been enough, and why keep the others?
You have finished the course. For the system these controls protect, see AI System Design; for what to watch while it runs, Observability Engineering.