# NVIDIA Can Sandbox the Agent Today. Sentry Is Still a Reference Design.

**Summary:** NVIDIA's OpenShell runtime is broadly available under Apache 2.0, while its BlueField-4 Sentry watchdog remains a reference design. I would test the software now and keep the hardware promise out of the security case until it can be deployed and challenged.

- Canonical: https://markhuang.ai/news/nvidia-openshell-available-sentry-reference-design
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-28
- Section: News
- Tags: NVIDIA, OpenShell, AI Agents, AI Security, Sandboxing
- Source: [MadRobot](https://madrobot.blog/2026/09/28/nvidia-open-agent-safety-platform-openshell-sentry-rogue-ai-agents/)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![An abstract AI agent sits inside two separate containment layers while an independent hardware watchdog guards its network path](https://cdn.markhuang.ai/news/nvidia-openshell-available-sentry-reference-design/hero.webp)

*OpenShell is the inner runtime boundary. Sentry is meant to keep watching from a separate hardware trust domain.*

NVIDIA's Open Agent Safety Platform sounds like one product. I would evaluate it as two offers at very different stages. The [MadRobot report](https://madrobot.blog/2026/09/28/nvidia-open-agent-safety-platform-openshell-sentry-rogue-ai-agents/) says OpenShell is available now as free, open-source software, while Sentry is a reference design for a watchdog on BlueField-4 data processing units. The report found no price or customer availability date for Sentry.

NVIDIA announced the combined platform on September 28, 2026, alongside more than 100 participating organizations. Its headline promise is attractive: OpenShell constrains an agent in software, then Sentry watches from outside the host and can quarantine an agent in milliseconds if it crosses a boundary. That last performance claim comes from NVIDIA. I have not found an independent evaluation that reproduces it.

That split decides how I would respond. OpenShell is the part I can evaluate today. Sentry may become the stronger backstop, but a reference architecture is not yet an operational control. I would pilot the software and leave the hardware promise out of the security case until I can buy it, deploy it, and test it under failure.

## OpenShell is already a concrete choice

[NVIDIA's OpenShell documentation](https://docs.nvidia.com/openshell/about/overview) describes an open-source runtime for fleets of autonomous agents. It places agents in sandboxes with kernel-level isolation and applies declarative policy to files, network access, processes, and provider credentials. The project is [published on GitHub under Apache 2.0](https://github.com/NVIDIA/OpenShell).

Agents need risky privileges to be useful. A coding agent may need to read a repository, run a package manager, call an inference endpoint, and push a change. A prompt can ask it to behave. A runtime policy can deny an undeclared file path or network destination even when the model decides the action would help.

The current quickstart is specific enough to inspect. It supports Linux, Apple Silicon Macs, and experimental Windows support through WSL 2, with Docker, Podman, or host virtualization underneath. NVIDIA's docs also say outbound connections are denied unless policy allows them, while credentials can be resolved for approved endpoints without handing the raw secret to the agent process.

> **Info:**
>
> I would start with one reversible agent task, deny network access by default, and review every permission the task needs. The first success metric is not whether the agent finishes. It is whether the policy blocks an action I deliberately put outside the job.

## Sentry is the ambitious half

[NVIDIA's announcement](https://nvidianews.nvidia.com/news/open-agent-safety-platform) calls Sentry an out-of-band watchdog built on its DOCA software and intended to run on BlueField-4 DPUs. The design puts monitoring and policy enforcement in an isolated hardware domain, separate from the CPU and software environment running the agent. NVIDIA says Sentry can inspect requests and responses, verify agent identity, provide attested telemetry, and stop an agent in milliseconds.

I like the systems principle. If the host or agent process is the thing under suspicion, the final stop control should not live inside it. That is the same lesson I took from the [AISI cyber evaluation that reached real internet systems](/news/cyber-eval-open-door): task instructions described a boundary, but the available network path defined what the agent could actually do.

Readiness is where I pause. NVIDIA repeatedly describes Sentry as a reference system design. Its release disclaimer also says many announced products and features remain at various stages and may arrive only on a "when-and-if-available" basis. The architecture can still be sound. A slide showing two layers, though, is not evidence that both layers protect a deployment today.

## The partner count does not answer the deployment question

NVIDIA names Anthropic, Microsoft, Salesforce, SAP, Scale AI, SpaceXAI, banks, infrastructure vendors, and robotics companies among more than 100 organizations working with the platform's technologies. The list signals interest, but it combines several kinds of involvement. NVIDIA describes some companies as integrating OpenShell, others as collaborating on shared safety work, and others as offering infrastructure support.

[WIRED reached the same awkward point](https://www.wired.com/story/nvidias-answer-to-rogue-agents-is-an-open-source-ai-security-system/) in its coverage: it was unclear whether the full partner list had adopted OpenShell or whether NVIDIA was describing a broader set of relationships. A launch roster is not a deployment inventory. I want to know which control is running, where it is enforced, what failure it has stopped, and who tested the result.

Public discussion noticed the practical split quickly. In [one LocalLLaMA thread](https://www.reddit.com/r/LocalLLaMA/comments/1ws9ydg/nvidia_shipped_openshell_an_open_source_sandbox/), the author planned to move local agents into OpenShell but skip Sentry because it requires BlueField hardware. One comment is not a market survey, but it captures the decision many smaller teams face. The open runtime can be tried on existing systems. The separate silicon layer belongs to a different procurement and operations conversation.

## What I would trust now

I would not dismiss the platform because its most interesting layer is early. OpenShell gives teams a real way to move permissions out of the prompt and into an inspectable runtime policy. That is useful on its own, and the Apache license makes the boundary open to inspection and modification.

I also would not credit Sentry with preventing a past breach on the strength of a counterfactual. The [AI Incidents record](https://ai-incident.org/risk-signals/nvidia-introduces-agent-safety-platform-with-separate-watchdog) notes that NVIDIA suggested the platform might have stopped earlier agent access to Hugging Face, but no published recreation independently establishes that result. A credible test would show the exact policy, attack path, telemetry, containment time, and behavior when the host itself is compromised.

I would test OpenShell now and verify its policies against deliberate escape attempts. Sentry stays on the future-defense list until an operator can obtain it and challenge it. NVIDIA is right to put agent controls outside the model. Now it needs deployment evidence for the hardware half of the pitch.
