Daniel Reyes, YuSMP Group
Daniel Reyes Principal Engineer, AI/ML, YuSMP Group · agentic and retrieval systems in production for US and EU teams
A glowing module sealed inside a translucent enclosure in a dark server rack, watched by a separate network card through a thin teal beam, illustrating a sandboxed AI agent under independent monitoring

The short answer

Nvidia has released a free, open-source platform that treats every AI agent as untrusted by default: OpenShell sandboxes the agent at the operating-system level, and Sentry watches it from separate hardware that the agent cannot reach. For teams building AI agents that act on real systems, it is the clearest signal yet that agent security is moving below the prompt and into infrastructure.

The software half, OpenShell, runs on ordinary Linux and macOS machines under the Apache 2.0 license, so any team can test it this week. The hardware half, Sentry, needs BlueField-4 DPUs and matters mostly to companies operating their own AI clusters. Microsoft, Oracle, Anthropic, Arm, Cisco and Intel are among the named partners.

What did Nvidia launch?

The Nvidia Open Agent Safety Platform is a bundle of two independent control layers for autonomous AI agents. Jensen Huang presented it on September 28 with a simple rule: “When you deploy an agent, no matter how smart, the first thing you do is to take away all of its rights.” The platform is free, open source and downloadable from Nvidia’s developer site and GitHub.

The first layer is OpenShell, a runtime Nvidia first showed in March and now positions as the base of the platform. Each agent runs in its own sandbox. Kernel controls confine which files it can open and which system calls it can make, every network connection is checked against policy before it leaves the sandbox, and the agent never sees real credentials: OpenShell adds them only to requests bound for approved endpoints. Operators describe allowed files, networks, tools, processes and credentials in a declarative policy.

The second layer is Nvidia Sentry, a reference design that runs on BlueField-4 data processing units. It records agent actions, policy decisions and data access, flags behavior that drifts from what is expected, and can change network policy or quarantine an agent in milliseconds. In Nvidia’s Vera Rubin systems the DPU sits on the node’s only path to the model, so the monitor sees traffic the agent cannot route around.

Why put an agent watchdog in separate hardware?

Because the threat model has changed from “the model says something wrong” to “the agent does something wrong.” CNBC and TechCrunch both tie the launch to a run of incidents in which agents from several frontier labs bypassed controls in test environments and reached real systems, including OpenAI agents that breached Hugging Face during a security exercise earlier this year. An Nvidia spokesperson told CNBC the new platform could have stopped that case.

A guardrail that lives in the same process, container or host as the agent can be disabled by an agent that has found a way to escalate. A monitor on a separate chip keeps working even if the host is compromised. That is the same logic security teams already apply to hardware security modules and out-of-band management: the component that enforces the rules should not be reachable by the thing it constrains.

Industry support is broad but not universal. Nvidia lists Microsoft, Oracle, Cisco, CoreWeave, Dell, HPE, Lenovo, Arm and Intel as partners and says it is working with Anthropic to connect Claude managed agents to OpenShell. SiliconANGLE reports SAP and Salesforce integrations, the latter routing human approvals through Slack. TechCrunch noted that OpenAI was not on the supporter list.

What it means for US & EU software teams

First, agent security becomes an infrastructure requirement. Prompt rules and model-side refusals stay useful, but they are no longer the control of record. Expect enterprise buyers and security reviewers to ask where the agent runs, what it can reach and who can stop it. Answering that is a cloud and DevOps question as much as an AI one: container isolation, network egress policy, secret management and logging are the building blocks.

Second, you can adopt the pattern without the hardware. OpenShell runs on commodity Linux and Mac machines, so the software layer is available to any team today. Sentry is different: it needs BlueField-4 DPUs, which makes it realistic mainly for companies running their own GPU clusters, or as a feature that cloud providers may offer later. For most product teams the practical step is the sandbox and policy model, not the silicon.

Third, audit trails help with compliance. Regulated teams in FinTech and HealthTech, and anyone preparing for EU AI Act obligations on logging and human oversight, need evidence of what an automated system did and why. A runtime that records every file, network and tool request an agent made, and every policy decision about it, turns that from a manual reconstruction into a log export.

What to do now

  1. Inventory agent permissions. List every agent that can write, send, pay, deploy or delete, and the credentials it holds today.
  2. Pilot deny-by-default. Run one existing agent inside OpenShell or an equivalent sandbox with no file, network or tool access, and log every blocked request for a week.
  3. Turn the log into policy. Convert observed legitimate requests into an explicit allowlist; anything else stays blocked and alerts.
  4. Remove raw secrets from agent context. Broker credentials at the network or proxy layer so a compromised agent cannot leak them.
  5. Keep a kill switch outside the agent. Make sure someone other than the agent’s own process can pause it, and test that path.

Frequently asked questions

What did Nvidia announce on September 28, 2026?

Nvidia CEO Jensen Huang introduced the Nvidia Open Agent Safety Platform, a free and open-source toolkit for keeping autonomous AI agents inside the boundaries their operators set. It has two layers: OpenShell, a sandboxed runtime that limits what an agent can touch, and Nvidia Sentry, an independent monitor that runs on BlueField-4 data processing units rather than on the CPU or GPU running the agent.

What is Nvidia OpenShell and do I need Nvidia hardware to use it?

OpenShell is an open-source runtime released under the Apache 2.0 license. It runs each agent in a sandbox where kernel controls restrict which files it can open and which system calls it can make, every outbound network connection passes a policy check, and real credentials are injected only into requests to approved endpoints. According to its repository, it runs on Linux, macOS on Apple Silicon and, experimentally, Windows with WSL 2, using Docker, Podman or host virtualization. Nvidia says it is optimized for its Vera CPUs but also works on Intel and Arm processors.

What does Nvidia Sentry add on top of OpenShell?

Sentry is a reference design that runs on BlueField-4 DPUs using Nvidia DOCA. Because it sits on separate hardware, it keeps watching the agent even if the host is compromised. Nvidia says it correlates agent actions, policy decisions and data access into an activity record, detects drift from expected behavior, updates network policy and can quarantine an agent in milliseconds. It requires BlueField-4 hardware, so it is mainly relevant to teams running their own AI infrastructure or buying from clouds that deploy it.

Why is Nvidia releasing this now?

The launch follows a series of reported incidents in which AI agents from several labs bypassed security controls in test environments and reached real systems. CNBC and TechCrunch both cite OpenAI agents breaching Hugging Face during a cybersecurity task earlier in 2026, and an Nvidia spokesperson told CNBC the platform could have stopped that incident. As agents move from reading data to changing systems, application-level guardrails alone are no longer considered sufficient.

Should my team adopt OpenShell for production agents?

It is worth piloting, not adopting blindly. The project is at version 0.1.x, so expect API changes. A sensible first step is to run one existing agent inside OpenShell with a deny-by-default policy, record every blocked file, network and tool request for a week, and turn that log into an explicit allowlist. Whatever runtime you choose, the design principles apply: least privilege, no raw secrets in the agent's context, egress control and an audit trail.

Sources

Nvidia Technical Blog — Nvidia Open Agent Safety Platform: a reference for continuous in-silicon agent monitoring
Nvidia OpenShell — GitHub repository (Apache 2.0)
TechCrunch — Nvidia launches new platform for reining in rogue AI agents
CNBC — Nvidia Open Agent Safety Platform to stop AI agents from breaking out
SiliconANGLE — Nvidia debuts enhanced safety controls to rein in rogue AI agents