Skip to content
Back to Insights
Agentic AIBy KE Engineering Team

Sandboxing Agents That Write Code

Sandboxing Agents That Write CodeAgentic AI cover for Sandboxing Agents That Write Codeattack pathAGENTIC AISandboxing Agents ThatWrite Code// ASSUME EXECUTION

Once an agent can write code and execute it, the security question changes shape. The question used to be what the model might say. Now it's what arbitrary code can do on this machine, with this network and these credentials. That is the oldest question in systems security, and the answer has not changed: assume the code is hostile and contain it.

The mistake teams make is sandboxing the model, adding guardrails on what the model will write. That's the same mistake as filtering prompts for injection (see Prompt Injection: Defense in Depth): filtering what goes into a model cannot stop it from producing harmful code, so the sandbox has to contain the output. The model will eventually write something you didn't anticipate. Sandbox the execution.

What the sandbox needs

A code-executing agent needs an environment that is:

Ephemeral. Created for one task, destroyed after. Nothing persists between runs unless explicitly exported through a controlled path. A compromised run can't leave anything behind for the next one.

Isolated. Its own filesystem, its own process namespace, no view of the host or of other runs. A container is the floor. For anything touching sensitive data, use stronger isolation: microVMs such as Firecracker, or kernel-level isolation such as gVisor or Kata Containers, because container escapes are a known category and the model is generating the attack surface.

Egress-controlled. Default deny. The sandbox reaches the specific endpoints the task needs and nothing else. An agent with a code tool and open egress is a data exfiltration path.

Credential-free. No secrets in the environment. If the code needs to call an API, it calls a broker inside the trust boundary that holds the credential and enforces policy on the call. The sandbox never sees the key.

Resource-bounded. CPU, memory, disk, wall clock, and process count are capped. An agent that writes an infinite loop or a fork bomb is a Tuesday, not an incident, if the limits are set.

The shape of it

┌─────────────────────────────────────────────────┐│  agent runtime (privileged)                      ││                                                  ││  ┌────────────┐   spawn    ┌──────────────────┐ ││  │ orchestrator│──────────▶│ sandbox (microVM) │ ││  │             │◀──────────│  no creds         │ ││  └─────┬──────┘  result    │  egress: allowlist│ ││        │                   │  ttl: 300s        │ ││        ▼                   └─────────┬────────┘ ││  ┌────────────┐                      │          ││  │ credential │◀─────────────────────┘          ││  │ broker     │   "call API X with these args"  ││  └────────────┘   broker holds key, enforces    ││                   policy, makes the call        │└─────────────────────────────────────────────────┘

The sandbox asks the broker to do things. The broker decides. The key never leaves the trust boundary.

Filesystem discipline

Mount the task's inputs read-only. Give the sandbox one writable scratch directory. Export results through an explicit path with a size limit and a type check. A sandbox that can write anywhere is a sandbox that can modify its own inputs mid-task, which is a confusing class of bug and an exploitable attack.

What the model sees of failures

When sandboxed code fails, the error text returns to the model as context. That is fine. What must not return is anything about the host: paths, environment variables, other processes, network topology. Scrub sandbox output before it re-enters the model's context, or the model learns about your infrastructure one stack trace at a time.

Egress is the part people skip

Every other control above is standard. Egress is where teams cut corners because "the agent needs to install packages." Fine: allowlist the package registry, through a proxy, with a log. Do not open the network because one use case needs one destination.

The cost of not doing this

An agent with a code tool, open egress, and credentials in its environment is a remote code execution vulnerability that you deployed on purpose. It will be found.