Pidoku

Sandboxing and Egress

Advanced 45 min Difficulty 4/5 Lesson 04 of 05

Prerequisites Agent Threats, Deployment and Infrastructure

The idea in one minute#

When an agent runs code or commands, treat it as an attacker’s code already running on your infrastructure, and ask what it can reach. The answer should be: its own disposable workspace, and a short list of network destinations — nothing else. Achieving that takes three boundaries, each of which must hold independently: an isolation boundary stronger than a container, a file-system and credential boundary that leaves nothing worth stealing inside, and an egress boundary that denies by default. Of the three, egress is the one that most often decides whether a successful hijack becomes a data breach.

Sandboxes and Tool Execution covers the design and capacity side; this lesson is the security configuration.

A picture#

flowchart TB
  subgraph TRUSTED["Trusted zone"]
    direction LR
    H[":i-bot: <b>Agent harness</b><br/><small>model calls, policy, credentials</small>"]
    BR[":vault: <b>Credential broker</b>"]
  end
  subgraph UNTRUSTED["Sandbox: assume hostile"]
    direction LR
    SH[":i-terminal: <b>Shell, interpreter,<br/>browser, local MCP</b>"]
    WS[(":i-hard-drive: <b>Task workspace only</b><br/><small>copy-on-write</small>")]
    SH --- WS
  end
  H -->|"commands in, output out<br/>(output is untrusted)"| SH
  UNTRUSTED --> B1[":gvisor: <b>Boundary 1: isolation</b><br/><small>gVisor / microVM, seccomp,<br/>no host mounts, resource limits</small>"]
  SH -->|"all traffic, including DNS"| PX[":envoyproxy: <b>Boundary 3: egress proxy</b><br/><small>allowlist, TLS inspection,<br/>credential injection</small>"]
  BR -.->|"injects short-lived token<br/>on allowed requests"| PX
  PX --> A1[":github: <b>Source host</b>"]
  PX --> A2[":python: <b>Package mirror</b>"]
  PX -.->|"deny + alert"| X[":i-ban: <b>Internet, metadata,<br/>internal network</b>"]
  class H compute
  class BR memory
  class SH warn
  class WS memory
  class B1 neutral
  class PX queue
  class A1,A2 io
  class X warn

How it really works#

Boundary 1 — isolation#

SettingWhy
gVisor or a microVM, not a plain containerA shared host kernel is too large an attack surface for hostile code
One sandbox per taskNo state, process or file survives from one user’s task into another’s
Non-root user; no privileged mode; minimal capabilitiesLimits what a kernel bug can be reached with
Seccomp profileBlocks system calls the workload never needs
No host mounts, no container-runtime socket, no device accessEach is a direct escape path
CPU, memory, process-count, disk and time limitsA fork bomb or a crypto-miner is contained and killed
Read-only base imageThe sandbox cannot persist changes to its own tooling
Dedicated node poolAn escape lands on a node with no credentials and no neighbours worth attacking
Destroyed at the endNothing to clean, nothing to carry over

Boundary 2 — files and credentials#

  • Mount only this task’s workspace: one repository checkout, one upload set.
  • No secrets inside. Not in environment variables, not in files, not in the image. A process can read all of them and so can an injected instruction.
  • Credentials are injected at the proxy. The code makes an unauthenticated request to an allowed host; the egress proxy adds the token. The sandbox never holds it.
  • Where a credential must exist inside — a git push, for example — it is scoped to one repository and one operation and expires in minutes.
  • Protect configuration in the workspace. A cloned repository may contain agent instruction files, MCP configuration, editor tasks, git hooks and CI definitions. These are untrusted content that some tools execute automatically. Do not auto-load them; make them read-only to the agent; require review before anything in them runs.
  • Output is untrusted. Whatever the sandbox prints goes back into the model’s context and is subject to the same input rail and taint rules as a web page.

Boundary 3 — egress#

Without egress control, isolation only protects you; the user’s data in the workspace can still be posted anywhere. Deny by default and build the allowlist per task type.

RuleDetail
Default deny, all protocolsIncluding DNS, ICMP and anything over non-standard ports
Allowlist by host name, through a proxyNetwork-level IP allowlists are too coarse for shared hosting and CDNs
DNS only via the proxy’s resolverA direct DNS query to <secret>.attacker.example is exfiltration
Block metadata endpoints and internal ranges169.254.169.254, link-local, private ranges, the cluster’s service network
Restrict methods and paths where possibleRead-only access to a package index is GET, not POST
Beware allowlisted hosts that accept uploadsA code host, a paste service or a storage bucket on the allowlist is itself an exfiltration channel. Scope to your organisation’s paths, or proxy through an internal mirror
Inspect and logEvery request and every denial is recorded with the task ID; denials alert
Size and rate limitsA large outbound body from a sandbox is unusual and worth stopping

The same policy applies to the agent runtime itself. Fetch and search tools called by the harness are egress too, and need a destination allowlist and the taint rule.

Tooling: Kubernetes NetworkPolicy for the coarse layer; Cilium for DNS-aware and identity-aware policy; an HTTP egress proxy — Envoy or a purpose-built one — for host allowlists, TLS handling and credential injection.

Defence in depth inside the sandbox#

A sandbox makes dangerous commands survivable, not harmless to the task. A second layer of command policy still pays for itself:

  • Deny patterns for commands with no legitimate use in the task: writing outside the workspace, touching credential paths, disabling tooling.
  • Approval for commands that affect things outside the sandbox: pushing, publishing, deploying.
  • Parse, do not pattern-match on strings. rm -rf / can be written a hundred ways; shell obfuscation defeats naive filters. Treat command filters as a convenience layer and the sandbox as the control.

Browsers and computer use#

An agent driving a browser or a desktop has the largest attack surface of any tool: every page is untrusted content, and the browser holds sessions.

  • A fresh browser profile per task, with no saved passwords and no logged-in sessions, unless the task is specifically to act in one account.
  • Never the user’s own everyday browser profile, which is logged into everything.
  • A domain allowlist for navigation when the task permits one.
  • Approval before entering credentials, making purchases, submitting forms with personal data, or changing account settings.
  • Screenshots and page text are untrusted input; visible text in an image can carry instructions.
  • Downloads go to the sandbox, are not executed, and are scanned if they leave it.

A browser agent logged into the user’s accounts holds all three legs of the trifecta at once. It is workable only with tight scoping and approval on consequential actions.

Local agents on developer machines#

Coding agents on a laptop cannot always use a remote microVM. The same three boundaries apply, with the operating system’s tools:

  • OS-level sandboxing of the agent’s command execution (Linux namespaces and seccomp, macOS sandbox profiles), restricting writes to the project directory.
  • A local proxy enforcing a network allowlist.
  • No access to ~/.ssh, cloud credential files, browser profiles or password stores.
  • Review before any project-supplied configuration, hook or MCP server runs.
  • Or: a dev container or remote workspace that makes the laptop a thin client.

Testing the sandbox#

Assume it leaks until shown otherwise. A recurring test job should, from inside a sandbox:

try to reach          an arbitrary internet host by name and by IP
                      the cloud metadata endpoint
                      another sandbox, a cluster service, the node
                      DNS directly, bypassing the proxy
try to read           environment variables for anything secret
                      files outside the workspace
try to persist        write to the base image; survive a restart
try to exhaust        CPU, memory, disk, process table

Every item should fail, and every failure should produce the expected alert. Run it after each change to the sandbox image, the runtime or the network policy.

Remember this#

  • Three independent boundaries: isolation, files and credentials, egress.
  • Nothing worth stealing is inside the sandbox; credentials are injected at the proxy.
  • Egress is deny-by-default, by host name, DNS included, with metadata and internal ranges blocked.
  • Allowlisted hosts that accept uploads are still exfiltration channels.
  • Workspace configuration files are untrusted and must not auto-execute.
  • Test the sandbox from the inside, continuously.

Try it#

  1. Write the egress allowlist for a coding agent working on one private repository. Which entries could be abused for exfiltration, and how would you narrow them?
  2. List everything an agent on your laptop could read today. What would you move out of reach?
  3. Add two checks to the sandbox test list that are specific to your environment.

Check yourself#

  1. Why is isolation without egress control insufficient?
  2. How can code call an authenticated API without holding the credential?
  3. Why are configuration files in a cloned repository a risk?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom