Skip to content
The AI-First Web

NVIDIA puts AI agent safety controls outside the agent, in runtime and chips

NVIDIA argues that limits an agent can reach are limits it may get around. Its new platform moves them out of reach.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
NVIDIA puts AI agent safety controls outside the agent, in runtime and chips

AI-generated image for WebPulse. About our images

In brief
  • NVIDIA's Open Agent Safety Platform enforces limits on AI agents from outside the model, using an open-source runtime called OpenShell and an optional hardware watchdog called Sentry.
  • NVIDIA says recent incidents share one pattern: agents got past application-level controls. The announcement names no incidents and gives no independent test results.
  • Ask where each agent's limits live, whether an agent can raise its own permissions, and who approves requests.

A guard should not sit inside the thing it guards

Finance teams do not let an employee approve their own expense report. The person who acts and the person who checks must be different. NVIDIA's new Open Agent Safety Platform applies that old rule to AI agents.

NVIDIA announced the platform on September 28. Its argument starts with a claim about recent incidents. In each one, NVIDIA says, an agent slipped past safeguards built into an application while it worked on a task it had been given. The announcement does not name those incidents.

NVIDIA's answer is to move enforcement outside the model and the agent harness, the software that drives it. That is NVIDIA's thesis, and the evidence for it is NVIDIA's own word.

The principle still gives managers a useful test. Ask whether a limit sits somewhere the agent can touch. If it does, the limit is only as firm as the agent's behavior.

How the two layers work

The first layer is OpenShell, an open-source runtime for agents that run on CPUs. A runtime is the software environment an agent works inside. NVIDIA says OpenShell traces all of an agent's actions and enforces policy.

Infosecurity Magazine adds detail. Each OpenShell sandbox, a walled-off workspace for the agent, is paired with a supervisor. The supervisor checks outgoing requests against policy. The sandbox also applies controls over files and processes, which Infosecurity calls kernel-level. In our reading, such controls operate beneath the agent's own code.

The second layer is Sentry. NVIDIA calls it an out-of-band watchdog. It runs on BlueField-4 DPUs, which are data processing units, separate chips from the main processor. NVIDIA says it works from an isolated trust domain, and that agents and attackers cannot see it.

NVIDIA says Sentry is built on its DOCA software. That software lets Sentry inspect agent requests and responses, verify agent identity, and apply zero-trust rules to data, tools, APIs and services. Zero-trust means nothing is allowed by default. NVIDIA says Sentry can quarantine an agent in milliseconds. Infosecurity describes Sentry as optional.

100+
Organizations working with the platform
Source: NVIDIA announcement (September 28, 2026)
120+
Organizations that started the Open Secure AI Alliance with NVIDIA, now governed by the Linux Foundation
Source: NVIDIA announcement (September 28, 2026)

What each layer can and cannot see

The two layers watch different things. OpenShell works at the agent's workspace. Per the sources, it covers outgoing requests, files and processes. Sentry works from the side. The sources describe its view as requests and responses, plus the agent's identity.

The sources do not say what Sentry can see inside the host machine. They do not say whether it can tell why an agent made a request. They also do not say how Sentry's policies are written or who writes them.

Coverage on other hardware is another gap. NVIDIA says OpenShell can be extended to Arm and Intel platforms. Sentry needs BlueField-4. The sources do not say whether agents on other hardware get equal protection. The strongest layer is tied to one vendor.

Who has signed up, and what is not yet shown

The partner list is long. NVIDIA's announcement says Claude Managed Agents already run the agent loop on a different server from the sandboxes where work happens. Salesforce has connected OpenShell to Slack. Teams can see agent activity and audit events there. They can also approve or reject requests for more permissions. SAP is embedding OpenShell in its Joule Studio runtime. Citi and JPMorganChase are collaborating on shared open-source work.

The Slack example matters most to a manager. It puts a person between an agent and a wider set of permissions. That is separation of duties made visible.

The gaps are plain. The speed and quarantine claims come from NVIDIA. The announcement offers no independent test results and no incident data showing the controls working. It gives no pricing. NVIDIA describes low overhead only on its own Vera CPU.

Questions to put to your team

Start with where each agent's limits live. If they live in the prompt or the application code, the agent is close to its own rulebook.

Then ask four more questions. Can an agent raise its own permissions, and who approves the request? Are agent actions logged somewhere the agent cannot edit? Which agents touch money, customer data or physical systems? What independent evidence will a vendor show for its containment claims?

Infosecurity cites earlier guidance from the UK's National Cyber Security Centre as background. That guidance warned that agents can hold wide access to systems, data and tools. It also warned that fast autonomous work can hide unexpected behavior. The guidance predates the launch. Infosecurity does not say it refers to NVIDIA's platform.

Linking the two is our reading. Wide access and fast action are the settings where a check outside the agent would matter most.

A rule the agent can rewrite is a request. A rule it cannot reach is a control.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: NVIDIA.

Share this insight