The idea in one minute#
A secure AI system is the ordinary architecture with six controls placed at six boundaries: identity at the front door, a policy check on every tool call, least-privilege credentials issued per task, isolation around execution, a deny-by-default exit to the network, and an audit trail across all of it. None is exotic. What makes the architecture secure is that every path from untrusted text to a real effect crosses at least one control that does not depend on the model behaving.
A picture#
flowchart LR
U[":i-user: <b>User</b>"] --> IDP[":keycloak: <b>Identity provider</b><br/><small>SSO, user token</small>"]
IDP --> GW[":envoyproxy: <b>AI gateway</b><br/><small>1 authenticate, limit,<br/>screen input and output</small>"]
GW --> AG[":i-bot: <b>Agent runtime</b><br/><small>holds no standing privilege</small>"]
AG --> PDP[":opa: <b>2 Policy check</b><br/><small>every tool call</small>"]
PDP --> TG[":agentgateway: <b>Tool gateway</b><br/><small>approved MCP servers only</small>"]
AG --> STS[":vault: <b>3 Token service</b><br/><small>per-task, scoped,<br/>on behalf of the user</small>"]
STS -.-> TG
TG --> SYS[":i-server: <b>Business systems</b><br/><small>enforce the user's permissions</small>"]
AG --> SB[":gvisor: <b>4 Sandbox</b><br/><small>code runs here</small>"]
SB --> EG[":cilium: <b>5 Egress control</b><br/><small>allowlist</small>"]
PDP -.->|"high-risk action"| HU[":i-hand: <b>Human approval</b>"]
AG -.-> AUD[(":i-scroll-text: <b>6 Audit trail</b><br/><small>who, on whose behalf, what, why</small>")]
TG -.-> AUD
GW -.-> AUD
class U neutral
class IDP,STS memory
class GW,PDP,TG queue
class AG compute
class SB,EG warn
class SYS io
class HU neutral
class AUD memoryHow it really works#
1 — the front door: identity, limits, screening#
Every request is authenticated at the gateway and attributed to a user, tenant and application. The gateway enforces token and spend limits, which doubles as protection against denial-of-wallet attacks. It also runs input and output screening: classifiers for known injection patterns, detection of secrets and personal data. These are useful and probabilistic; they lower the rate of attacks and catch accidents, and the design must not depend on them.
2 — the policy check on every tool call#
Between the model’s proposal and its execution sits a policy decision made by code. The question it answers has five parts:
may this agent (identity of the agent and its version)
for this user (whose authority is being used)
call this tool (is it in the task's allowed set?)
with these arguments (recipient in allowlist? path inside workspace? amount under limit?)
given this session's taint (has untrusted content entered the context?)The last line is what makes the rule of two enforceable. The runtime tracks whether the session has ingested untrusted content; once it has, tools that send data out or make irreversible changes require approval or are denied. Policies are written in a policy language or in code, live in version control, and are tested like any other code.
3 — credentials: per task, scoped, on behalf of#
The agent has no API keys of its own lying around. When a tool call is approved, the runtime obtains a credential for that call:
- On behalf of the user. The user’s identity is exchanged for a token that carries both the user and the agent, so the business system applies the user’s permissions and the audit log shows who really acted.
- Scoped down. Narrower than the user’s own access: this repository, this mailbox, read only.
- Short-lived. Minutes. A leaked token expires before it is useful.
- Bound to the audience. A token issued for one MCP server is not accepted by another.
The MCP authorization model (OAuth 2.1, with enterprise-managed authorization through the organisation’s identity provider) and workload-identity systems such as SPIFFE are the building blocks; Identity and Authorization shows how they fit together.
4 — isolation around execution#
Model-written code and local tools run in a per-task sandbox with a user-space kernel or microVM boundary, a workspace containing only that task’s files, and no secrets. The agent runtime, which talks to the model and holds session state, runs outside it. See Sandboxes and Tool Execution.
5 — egress: deny by default#
All outbound traffic from sandboxes and from the agent runtime passes a proxy or network policy with an allowlist. This is the control that still holds when everything above has been bypassed: stolen data needs a way out. The same principle applies to what the user’s browser does with model output — rendered images and links are restricted to trusted hosts.
6 — the audit trail#
One record per action, joined by a trace ID:
when timestamp
who agent identity and version; the user it acted for; the tenant
what tool, arguments (redacted where needed), result status
why the task; the model call that proposed it; policy decision and rule
with what credential scope used; sandbox ID
provenance which untrusted sources were in the context at that pointProvenance is what lets an investigator answer “which document caused this?” It also drives detection: an agent that suddenly calls a tool it has never used, or sends to a new domain, after reading external content, is the pattern to alert on.
The supply side#
The controls above protect the running system. The same architecture needs a supply chain that can be trusted:
| Component | Control |
|---|---|
| Model weights | Pulled from your registry only; safe serialisation format; scanned; signed |
| MCP servers and tools | An approved catalogue; pinned versions; descriptions reviewed; re-reviewed on change |
| Prompts, policies, routing | In version control, reviewed, evaluated before release |
| Agent frameworks and dependencies | Ordinary software supply-chain hygiene: lockfiles, scanning, provenance |
| Retrieval corpus | Known writers; ingestion scanning; the ability to remove a poisoned document |
Mapping to the recognised frameworks#
A reviewer will ask how the design addresses the standard lists. The short version:
| Control | OWASP LLM Top 10 (2026) | OWASP Agentic Top 10 |
|---|---|---|
| Policy layer, taint tracking | Prompt Injection, Excessive Agency | Agent Goal Hijack, Tool Misuse |
| Per-task scoped credentials | Excessive Agency, Sensitive Information Disclosure | Identity and Privilege Abuse |
| Sandbox and egress control | Improper Output Handling | Unexpected Code Execution |
| Approved tool catalogue, signed models | Supply Chain, Data and Model Poisoning | Agentic Supply Chain Vulnerabilities |
| Guarded memory writes | Vector and Embedding Weaknesses | Memory and Context Poisoning |
| Token limits and budgets | Unbounded Consumption | Cascading Failures |
| Approval on high-risk actions | Excessive Agency | Human-Agent Trust Exploitation |
| Audit trail, behaviour monitoring, kill switch | — | Rogue Agents |
Frameworks and Maps explains each list.
What it costs#
Security controls are not free, and a design should state the price:
| Control | Latency | Operational cost |
|---|---|---|
| Gateway screening | 20–200 ms per call, if a classifier model runs | A small model to serve |
| Policy check | Under 5 ms | Policies to write and test |
| Token exchange | 10–50 ms, cacheable per task | An identity service |
| Sandbox | 100–300 ms at task start | A node pool and controller |
| Egress proxy | 1–5 ms | Allowlists to maintain |
| Human approval | Seconds to hours | People’s attention — the scarcest resource |
The expensive one is human approval. Spend it only on actions that are irreversible, external or high-value, and make every other path safe by construction.
Review checklist#
- Every source of context text is listed with who controls it.
- No agent context holds private data, untrusted content and an outbound channel without an approval gate.
- Every tool call passes a policy check written in code.
- Tools act with the user’s authority, scoped down, with short-lived credentials.
- Model-written code runs in an isolated sandbox with no secrets.
- Egress is deny-by-default for sandboxes and for the agent runtime.
- Rendered output cannot load content from untrusted hosts.
- Memory writes from untrusted content are blocked or reviewed.
- Models, tools and MCP servers come from an approved, pinned catalogue.
- Every action is logged with agent, user, arguments and context provenance.
- A task, an agent and a tenant can each be stopped immediately.
- Adversarial cases are in the evaluation suite and run on every release.
Remember this#
- Six controls at six boundaries: front door, policy check, scoped credentials, sandbox, egress, audit.
- The policy check sees the session’s taint; that is how the rule of two is enforced.
- Agents act on behalf of users, never with standing service privileges.
- Egress control is the last line and holds when the others fail.
- State the cost of each control; spend human approval sparingly.
Try it#
- Take the worked design you care most about and run the checklist. List the unchecked items in order of blast radius.
- Write the policy, in plain sentences, for a
send_emailtool in a session that has read an external web page. - Write the audit record for one tool call in a system you know. Which fields are missing today?
Check yourself#
- Which of the six controls still protects data after the model has been fully manipulated?
- Why should an agent use the user’s delegated authority rather than a service account?
- Why is provenance recorded in the audit trail?