Skip to content
Innovation & Growth

'Sandboxed' tells you little until you ask what the sandbox leaves shared

Endor Labs' primer says the label hides what is shared. For AI agents, it points to action policy and logs.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
'Sandboxed' tells you little until you ask what the sandbox leaves shared

AI-generated image for WebPulse. About our images

In brief
  • Endor Labs says "sandboxed" means little until you know which resources a sandbox isolates, which it shares and what happens if it is bypassed.
  • The primer says AI agents hold authority you granted, so prompt injection can steer them. It points to action policy, narrow credentials and logs. We argue someone must own them.
  • Replace "is it sandboxed?" with questions on shared resources, metadata access, outbound limits and session logs, and give each to a named owner.

A word that answers nothing

When a vendor or a questionnaire says "sandboxed", the box is usually ticked and the discussion ends. Endor Labs, in a primer published October 2, 2026, argues that the word on its own means nothing.

The reason is simple. Every sandbox isolates some resources and leaves others shared. The primer says the interesting incidents happen in the gap between what was locked down and what people assumed was.

For AI agents, the primer says the boundary moves from what code can touch to what the agent may do. It points to an action policy and to logging as the controls that carry that weight.

The primer does not discuss who runs those controls. That part is our argument. In our reading, someone has to be named, trained and paid to run the policy and the logs. Otherwise "sandboxed" is replaced by another label nobody checks.

Three jobs, so three different claims

The researchers say sandboxes do one of three jobs: stop hostile code, limit the damage from buggy code, or make test runs repeatable. A boundary built for one job can fail at another. A vendor who says "sandboxed" has not told you which job the boundary was built to do.

Strength costs money, and setup decides it

A virtual machine gives code its own operating system kernel, enforced by a hypervisor. It costs memory and start-up time. Firecracker, a lightweight option, boots one in around 125 milliseconds.

~125 ms
Firecracker microVM boot time
Source: Endor Labs, sandbox primer (October 2, 2026)

Even that boundary is not absolute. The primer cites VENOM (CVE-2015-3456), where guest code escaped through QEMU's virtual floppy disk controller.

Containers share one kernel with the host. That makes them fast, and a kernel bug becomes a shared failure. The primer notes that runc, the tool that launches containers, has produced several escape bugs. One of them, CVE-2019-5736, let a hostile container image tamper with runc on the host and gain code execution there.

The researchers say the technology is not the variable. The setup is. A container run in privileged mode, or with the Docker socket mounted, is weak. Harden the same container and the primer calls it a serious boundary. That means filtering system calls tightly, making the root filesystem read-only, running as a non-root user and blocking outbound traffic. Two suppliers can both say "containers" and deliver very different protection.

Two holes the label hides

The first is the network. Code cut off from your files may still reach the cloud's internal address, 169.254.169.254. The primer says that can often let it obtain credentials for the host's cloud role.

The second is identity. Processes inherit environment variables, credential files and SSH agent sockets. The filesystem can be sealed completely while the code still holds a token that works elsewhere.

Agents move the boundary into policy

An AI coding agent is useful because it can read your repository, run tests, install dependencies and open pull requests. Seal it off fully and it cannot do that work.

So the primer says the question changes from what a process can touch to what the agent may do, with which credentials, and who confirms. It calls the agent a confused deputy. It is not malware. It holds authority you gave it, and someone else may steer it through instructions hidden in an issue, a README or a web page.

The primer borrows Simon Willison's "lethal trifecta". An attack needs private data, untrusted content and a route to the outside. Take away any one of the three and the attack loses its value.

3
Ingredients in the "lethal trifecta"
Source: Simon Willison's framing, as cited by Endor Labs (October 2, 2026)

The primer says serious setups run the agent inside a container or microVM, then add an action policy on top. The container handles a malicious build script. The policy handles an agent talked into something that looks reasonable and is wrong.

The primer describes what that policy covers. It sets which directory the agent sees, which destinations it may reach and which actions need a human. It calls limiting outbound access the highest-value control and the one most often skipped. It also warns that constant approval prompts wear people down, so they stop reading what they approve.

Logging is the other half. An agent's permissions build up over a session, so a static settings file cannot show what it ended up allowed to do. Only records made at the time can. The primer points to commands run, files changed, tools called and who signed off.

Who owns each question

The assignments below are our suggestion. The primer supplies the questions, not the owners.

Vendor risk or procurement: ask which resources the sandbox isolates, which it shares and what happens if it is bypassed. An acceptable answer names the boundary type and lists what is shared. The word alone is not an answer.

Cloud platform owner: ask whether sandboxed code can reach the metadata address and what credentials it inherits. An acceptable answer is no to the first. For the second, credentials should be short-lived and limited to one repository.

Engineering lead: ask whether coding agents run in a container or microVM with outbound access limited to your package registry and Git host. An acceptable answer is a written allowlist, not a verbal assurance.

Security or audit lead: ask what the logs would show if an agent were talked into something wrong. An acceptable answer is a session record that lets you rebuild commands, changes and approvals.

Finance and the executive sponsor: ask who owns the agent policy and whether their time is funded. An acceptable answer is a named person, not "the platform".

The primer says prompt injection has no clean technical fix. Plan for a successful injection. Decide now what it can see, where it can send it, and which steps need a person.

A boundary you have not described is a boundary you have assumed.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Endor Labs.

CVEs in this analysis
CVE-2019-5736 CVE-2015-3456
Share this insight