NVIDIA proposes an open safety layer from agent runtime to hardware

NVIDIA’s Open Agent Safety Platform combines OpenShell on Vera CPUs with Sentry on BlueField-4 DPUs to provide out-of-band monitoring and enforcement for long-running agents.

NVIDIA introduced the Open Agent Safety Platform on September 28, 2026, as a reference architecture that extends agent safety from the application layer into runtime and hardware. It places NVIDIA OpenShell beside workloads on Vera CPUs and NVIDIA Sentry on BlueField-4 DPUs for independent monitoring and enforcement, with the goal of preventing an agent from governing its own behavior when it drifts from the intended task.

NVIDIA describes five design principles: policies should be verifiable, enforcement should be out of band, the path to the model should be a control point, agent authority should scale with the ability to observe behavior, and labs, enterprises, and hardware providers should share responsibility. This differs from putting safety instructions only in a system prompt because the main control points live in runtime or infrastructure layers that the agent is not supposed to modify.

The architecture has three layers. The application contains models, harnesses, tools, and data; the runtime projects the workload onto a workstation, edge device, or data center while providing monitoring and policy enforcement; and the infrastructure includes networks, databases, filesystems, general compute, and accelerated safety hardware. OpenShell provides sandboxing, identity, permissions, and policies, while Sentry on BlueField-4 provides an independent activity record and control path.

NVIDIA says OpenShell can turn an agent policy into verifiable boundaries around files, networks, tools, processes, and credentials. Using NVIDIA DOCA, Sentry correlates agent interactions, policy decisions, and tool and data access into a contextual record for detecting drift and tracking delegated authority. These are NVIDIA’s architectural claims and still need validation against real hardware, drivers, identity systems, and policy maintenance costs.

For Vera Rubin POD deployments, NVIDIA places BlueField-4 on the node’s path to the model so that it can act as an out-of-band observation point that is difficult for an agent to bypass. The article also says the platform is compatible with other hardware, but it does not provide an independent benchmark or complete production failure data here. A reference design should therefore not be treated as an immediately deployable answer for every enterprise.

The practical value of this direction is turning ‘what an agent should not do’ into runtime policy that can be tested, logged, and blocked. Hardware monitoring does not automatically solve faulty policies, excessive authority, data classification, supply-chain risk, or human approval. Organizations still need threat models, least privilege, rollback design, audit retention, end-to-end tests, and clear ownership for changing or approving policies.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.