The idea in one minute#
When an agent runs code or commands, treat it as an attacker’s code already running on your infrastructure, and ask what it can reach. The answer should be: its own disposable workspace, and a short list of network destinations — nothing else. Achieving that takes three boundaries, each of which must hold independently: an isolation boundary stronger than a container, a file-system and credential boundary that leaves nothing worth stealing inside, and an egress boundary that denies by default. Of the three, egress is the one that most often decides whether a successful hijack becomes a data breach.
Sandboxes and Tool Execution covers the design and capacity side; this lesson is the security configuration.
A picture#
flowchart TB
subgraph TRUSTED["Trusted zone"]
direction LR
H[":i-bot: <b>Agent harness</b><br/><small>model calls, policy, credentials</small>"]
BR[":vault: <b>Credential broker</b>"]
end
subgraph UNTRUSTED["Sandbox: assume hostile"]
direction LR
SH[":i-terminal: <b>Shell, interpreter,<br/>browser, local MCP</b>"]
WS[(":i-hard-drive: <b>Task workspace only</b><br/><small>copy-on-write</small>")]
SH --- WS
end
H -->|"commands in, output out<br/>(output is untrusted)"| SH
UNTRUSTED --> B1[":gvisor: <b>Boundary 1: isolation</b><br/><small>gVisor / microVM, seccomp,<br/>no host mounts, resource limits</small>"]
SH -->|"all traffic, including DNS"| PX[":envoyproxy: <b>Boundary 3: egress proxy</b><br/><small>allowlist, TLS inspection,<br/>credential injection</small>"]
BR -.->|"injects short-lived token<br/>on allowed requests"| PX
PX --> A1[":github: <b>Source host</b>"]
PX --> A2[":python: <b>Package mirror</b>"]
PX -.->|"deny + alert"| X[":i-ban: <b>Internet, metadata,<br/>internal network</b>"]
class H compute
class BR memory
class SH warn
class WS memory
class B1 neutral
class PX queue
class A1,A2 io
class X warnHow it really works#
Boundary 1 — isolation#
| Setting | Why |
|---|---|
| gVisor or a microVM, not a plain container | A shared host kernel is too large an attack surface for hostile code |
| One sandbox per task | No state, process or file survives from one user’s task into another’s |
| Non-root user; no privileged mode; minimal capabilities | Limits what a kernel bug can be reached with |
| Seccomp profile | Blocks system calls the workload never needs |
| No host mounts, no container-runtime socket, no device access | Each is a direct escape path |
| CPU, memory, process-count, disk and time limits | A fork bomb or a crypto-miner is contained and killed |
| Read-only base image | The sandbox cannot persist changes to its own tooling |
| Dedicated node pool | An escape lands on a node with no credentials and no neighbours worth attacking |
| Destroyed at the end | Nothing to clean, nothing to carry over |
Boundary 2 — files and credentials#
- Mount only this task’s workspace: one repository checkout, one upload set.
- No secrets inside. Not in environment variables, not in files, not in the image. A process can read all of them and so can an injected instruction.
- Credentials are injected at the proxy. The code makes an unauthenticated request to an allowed host; the egress proxy adds the token. The sandbox never holds it.
- Where a credential must exist inside — a git push, for example — it is scoped to one repository and one operation and expires in minutes.
- Protect configuration in the workspace. A cloned repository may contain agent instruction files, MCP configuration, editor tasks, git hooks and CI definitions. These are untrusted content that some tools execute automatically. Do not auto-load them; make them read-only to the agent; require review before anything in them runs.
- Output is untrusted. Whatever the sandbox prints goes back into the model’s context and is subject to the same input rail and taint rules as a web page.
Boundary 3 — egress#
Without egress control, isolation only protects you; the user’s data in the workspace can still be posted anywhere. Deny by default and build the allowlist per task type.
| Rule | Detail |
|---|---|
| Default deny, all protocols | Including DNS, ICMP and anything over non-standard ports |
| Allowlist by host name, through a proxy | Network-level IP allowlists are too coarse for shared hosting and CDNs |
| DNS only via the proxy’s resolver | A direct DNS query to <secret>.attacker.example is exfiltration |
| Block metadata endpoints and internal ranges | 169.254.169.254, link-local, private ranges, the cluster’s service network |
| Restrict methods and paths where possible | Read-only access to a package index is GET, not POST |
| Beware allowlisted hosts that accept uploads | A code host, a paste service or a storage bucket on the allowlist is itself an exfiltration channel. Scope to your organisation’s paths, or proxy through an internal mirror |
| Inspect and log | Every request and every denial is recorded with the task ID; denials alert |
| Size and rate limits | A large outbound body from a sandbox is unusual and worth stopping |
The same policy applies to the agent runtime itself. Fetch and search tools called by the harness are egress too, and need a destination allowlist and the taint rule.
Tooling: Kubernetes NetworkPolicy for the coarse layer; Cilium for DNS-aware and identity-aware policy; an HTTP egress proxy — Envoy or a purpose-built one — for host allowlists, TLS handling and credential injection.
Defence in depth inside the sandbox#
A sandbox makes dangerous commands survivable, not harmless to the task. A second layer of command policy still pays for itself:
- Deny patterns for commands with no legitimate use in the task: writing outside the workspace, touching credential paths, disabling tooling.
- Approval for commands that affect things outside the sandbox: pushing, publishing, deploying.
- Parse, do not pattern-match on strings.
rm -rf /can be written a hundred ways; shell obfuscation defeats naive filters. Treat command filters as a convenience layer and the sandbox as the control.
Browsers and computer use#
An agent driving a browser or a desktop has the largest attack surface of any tool: every page is untrusted content, and the browser holds sessions.
- A fresh browser profile per task, with no saved passwords and no logged-in sessions, unless the task is specifically to act in one account.
- Never the user’s own everyday browser profile, which is logged into everything.
- A domain allowlist for navigation when the task permits one.
- Approval before entering credentials, making purchases, submitting forms with personal data, or changing account settings.
- Screenshots and page text are untrusted input; visible text in an image can carry instructions.
- Downloads go to the sandbox, are not executed, and are scanned if they leave it.
A browser agent logged into the user’s accounts holds all three legs of the trifecta at once. It is workable only with tight scoping and approval on consequential actions.
Local agents on developer machines#
Coding agents on a laptop cannot always use a remote microVM. The same three boundaries apply, with the operating system’s tools:
- OS-level sandboxing of the agent’s command execution (Linux namespaces and seccomp, macOS sandbox profiles), restricting writes to the project directory.
- A local proxy enforcing a network allowlist.
- No access to
~/.ssh, cloud credential files, browser profiles or password stores. - Review before any project-supplied configuration, hook or MCP server runs.
- Or: a dev container or remote workspace that makes the laptop a thin client.
Testing the sandbox#
Assume it leaks until shown otherwise. A recurring test job should, from inside a sandbox:
try to reach an arbitrary internet host by name and by IP
the cloud metadata endpoint
another sandbox, a cluster service, the node
DNS directly, bypassing the proxy
try to read environment variables for anything secret
files outside the workspace
try to persist write to the base image; survive a restart
try to exhaust CPU, memory, disk, process tableEvery item should fail, and every failure should produce the expected alert. Run it after each change to the sandbox image, the runtime or the network policy.
Remember this#
- Three independent boundaries: isolation, files and credentials, egress.
- Nothing worth stealing is inside the sandbox; credentials are injected at the proxy.
- Egress is deny-by-default, by host name, DNS included, with metadata and internal ranges blocked.
- Allowlisted hosts that accept uploads are still exfiltration channels.
- Workspace configuration files are untrusted and must not auto-execute.
- Test the sandbox from the inside, continuously.
Try it#
- Write the egress allowlist for a coding agent working on one private repository. Which entries could be abused for exfiltration, and how would you narrow them?
- List everything an agent on your laptop could read today. What would you move out of reach?
- Add two checks to the sandbox test list that are specific to your environment.
Check yourself#
- Why is isolation without egress control insufficient?
- How can code call an authenticated API without holding the credential?
- Why are configuration files in a cloned repository a risk?