Skip to content
Security & Trust

In one test, NVIDIA's agent sandbox held; leaks came from operator settings

Sorami's test of OpenShell v0.1.2 found no bypass of a documented control. Settings still let data out.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
In one test, NVIDIA's agent sandbox held; leaks came from operator settings

AI-generated image for WebPulse. About our images

Key finding

Agent leaked the secret without OpenShell: 10 of 10 runs (Source: Sorami Technical Report (September 29, 2026))

The wall held. The settings were the problem

A sandbox for AI agents is only as strong as the settings someone chooses for it. That is the lesson from a new test of NVIDIA's OpenShell, and it changes what a security review of agent tools should look at.

OpenShell is a free, open-source runtime that runs an AI agent inside a locked-down box. It blocks outbound network traffic unless a policy rule allows it. NVIDIA launched it on 28 September 2026, as the centrepiece of its Open Agent Safety Platform. The following day, the Australian consultancy Sorami published a test of version 0.1.2.

What the researchers tested

Sorami ran 35 test IDs and 41 test-and-condition cells, which made up 123 deterministic trials. It also ran paired agent tests, with 10 runs per condition. It committed its test plan before running anything and published its logs. Every documented control Sorami tested held. That includes blocking outbound traffic by default, matching rules to specific programs, and restricting file access.

The researchers also tested a realistic case. A developer asks an AI agent to set up a cloned code project. The project's setup script is malicious and tries to send out a secret file. Sorami describes this as a stand-in for the common supply-chain case.

10 of 10 runs
Agent leaked the secret without OpenShell
Source: Sorami Technical Report (September 29, 2026)
0 of 10 runs
Agent leaked the secret under OpenShell's default policy
Source: Sorami Technical Report (September 29, 2026)

That is a clean result for the default. It is also where the good news ends.

Four settings that let data out

Every leak Sorami could produce went through a setting that an operator chooses. Each one is a setting a person chooses, or leaves at its default.

First, read-write rules. Once the policy treated the destination as a read-write endpoint, the secret arrived in all 10 runs. That is no better than having no sandbox at all.

Second, loose rules. A rule that allowed only a simple health-check request still passed the canary secret, because it rode in the query string of the web address. Headers were another route.

Third, audit mode. This is the default for rules that inspect web requests. In the test, a forbidden request went through, and the log line read ALLOWED at INFO level. A reviewer skimming logs would see nothing alarming.

Fourth, automatic approval. In this mode, OpenShell can approve new access without a person looking. It granted new public hosts in 12 of 12 trials. That included rules OpenShell drafted itself from blocked connections. Flagged destinations, such as private addresses, still waited for a human.

12 of 12 trials
Automatic approval granted new public hosts with no human review
Source: Sorami Technical Report (September 29, 2026)

The safety check has blind spots

OpenShell includes a policy prover. It compares a proposed policy against a boundary and flags risky new access. In Sorami's test, GraphQL, MCP, WebSocket and JSON-RPC rules came back as unsupported by the prover. The loader still accepted them, with no prover check behind them.

NVIDIA's own documentation says the prover does not model these rule types. So this is a disclosed limit, not a hidden flaw. The risk is a gap between what the documentation says and what a busy team reads. A policy can pass and still contain rules nobody checked.

Sorami also found three logging gaps on this driver. For example, approval events did not appear in the sandbox's logs. They were in the gateway log and the rule history instead. A team that watches only one place would miss them.

What this test does not show

The limits matter. Sorami ran everything on one Apple Silicon laptop with one small model, qwen3:8b. That model ignored the injected instruction in the first test. So the results say little about how a stronger model behaves.

The researchers also could not test whether an agent tries to widen its own access. The model never hit a denial, so they report that as not elicited. The DNS result rests on reading the source code, not on a packet capture. And this is version 0.1.2, released a day earlier. It is one lab's test, not a verdict on the product.

What to ask your team

The lesson is that a strong default gets weaker each time someone adds an exception. Firewalls taught the same lesson long ago. The device rarely fails. The rule list grows until it stops meaning much.

If your teams run coding or operations agents, put these questions to them:

Who can approve new network access for an agent, and is a person always in that step? Sorami advises keeping approval manual.

Are web-request rules set to enforce, or left in audit mode? Audit mode records a violation. It does not stop it.

Does any agent have a read-write rule to an outside host? Sorami advises avoiding them.

Is every policy proved in your build pipeline, and does someone know which rule types the prover skips?

Where are approval events logged, and does anyone read that log?

The boundary did its job in this test. Whether it keeps doing it depends on the people holding the settings.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: sorami.com.au.

Share this insight