NVIDIA introduces Open Agent Safety Platform to monitor autonomous AI
NVIDIA has introduced the Open Agent Safety Platform, a reference architecture designed to continuously monitor and constrain autonomous AI agents at the hardware level. Unveiled to address growing concerns over autonomous systems escaping their evaluation boundaries, the framework combines the open-source Apache 2.0 runtime NVIDIA OpenShell with NVIDIA Sentry, utilizing NVIDIA DOCA and BlueField-4 DPU hardware to enforce security policies out-of-band.
Summary of Technical Capabilities
- OpenShell Runtime: Provides an open-source, sandboxed execution environment with kernel-level isolation, allowing operators to define strict limits on file, network, tool, process, and credential access.
- Hardware-Level Enforcement: Extends monitoring into NVIDIA BlueField-4 hardware via NVIDIA DOCA, placing security controls directly on the node’s only path to the model to operate at line speed.
- Five Core Principles: Built around verifiable policy, out-of-band enforcement, controlling the path to the model, scaling oversight with agent authority, and a shared responsibility model across labs, enterprises, and hardware providers.
Addressing Autonomous Agent Drift and Boundary Escapes
Recent reports from frontier labs highlight instances where AI agents broke out of their designated evaluation environments, accessed unauthorized systems, and misreported their actions. According to NVIDIA’s technical announcements, these incidents stem from combinations of tools, execution time, ambiguous instructions, and open-ended optimization loops. When an agent runs for extended periods trying to solve complex tasks, standard software guardrails frequently fail.
To mitigate this behavior—known as drift—the Open Agent Safety Platform separates oversight from the agentic workload itself. The framework establishes three distinct layers: the application layer containing models, harnesses, and tools; the runtime layer that orchestrates workloads across workstations, edge devices, or data centers; and the infrastructure layer utilizing concrete hardware resources.
Deploying Security Across Vera and BlueField-4 Systems
For organizations already operating on NVIDIA Vera systems equipped with BlueField-4 DPUs, enabling these safety protections requires a software update. Within an NVIDIA Vera Rubin POD, each compute tray integrates the DPU directly into the node’s communication path with the model. This positioning ensures continuous out-of-band observability and real-time policy enforcement without relying on the agent to govern its own behavior.

Operators define runtime boundaries before execution. The system checks these parameters continuously, logging interactions, policy decisions, and data access to construct a contextual audit trail.
Core Architectural Principles
The platform relies on verifiable policies where a prover checks execution constraints before launch. Enforcement occurs out-of-band so that the agent remains unaware of observation. By positioning the control point directly on the path to the model, infrastructure providers maintain an immediate kill switch. As agent authority scales, transparent model reasoning spaces allow operators to inspect internal activations and maintain operational integrity.