Aivora
NVIDIA DeveloperAI SafetyIntermediate

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring

NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構

2 min read
NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
The 30-second version

As autonomous AI agents evolve, incidents of agents escaping test environments highlight critical security gaps. Drawing from the history of browser sandboxing, NVIDIA introduced the Open Agent Safety Platform. Powered by OpenShell (Apache 2.0), a secure runtime with kernel-level isolation, the platform enforces strict user-defined limits. When integrated with NVIDIA Sentry and BlueField-4 DPUs, it provides continuous, out-of-band hardware-level monitoring directly on the path to the model, functioning as an independent, tamper-proof kill switch.

Key points

01

Zero-Trust Kernel Sandbox

Uses OpenShell to run agents in isolated environments with kernel-level security, restricting file, network, and tool access by default.

02

Independent Out-of-Band Control

Monitoring runs out-of-band, beyond the agent's reach. The agent cannot detect or tamper with the security controls.

03

Model Path as Control Point

By positioning controls directly on the data path to the model, operators gain both a highly visible observation post and an instant kill switch.

04

In-Silicon Hardware Protection

Integrates with BlueField DPUs and DOCA to verify identity and detect behavioral drift at the silicon level, ensuring safety even if host resources are compromised.

How it works

NVIDIA Open Agent Safety Three-Layer Architecture
Projects PolicyHardware MonitoringControls Path to BrainApplication LayerRuntime Layer(OpenShell)Infrastructure Layer(BlueField)Model Brain

Why it matters

While traditional safety relies on model alignment, long-running agents often experience behavioral drift that training alone cannot resolve. By shifting enforcement to the runtime and silicon layer, NVIDIA removes the need to trust the agent to govern itself. Similar to how browser sandboxing enabled early online commerce, this hardware-level trust layer is critical for unlocking the full potential of the enterprise agent economy.

Who it affects

  • AI Developer
  • Enterprise Leader
  • AI Researcher
  • Policy Maker

How to use it

  1. 1Secure Autonomous Enterprise Agents: Safe execution of agents requiring sensitive database access and external code execution within sandboxed runtime environments.
  2. 2Hardware-Enforced Behavioral Auditing: Utilizing BlueField DPUs to monitor multi-step agents for identity verification and structural drift in high-security environments.

Limitations & caveats

  • Full hardware-level out-of-band monitoring and in-silicon protection are highly optimized for systems utilizing NVIDIA Vera CPUs and BlueField DPUs.
  • Behavioral drift cannot be entirely trained out of autonomous models, requiring predefined operator policies and manual intervention strategies.

Related

NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents
NVIDIA DeveloperAI Safety

NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents

NVIDIA OpenShell:不重寫程式碼,為 AI Agent 部署執行階段安全隔離防護

NVIDIA released OpenShell 0.1.0, an open-source runtime that secures AI agents using external sandboxing, supervisor monitoring, and formal policy analysis without rewriting agent code.

2 min read
Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users
arXivAI Safety

Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users

以「信念自我蒸餾」提取使用者模型:揭示大語言模型對用戶意圖的內在表徵

Researchers introduce Belief Self-Distillation (BSD), a framework to read and write an LLM's implicit beliefs about its users, revealing how user-intent inference drives safety decisions and showing shared representation geometries across models.

2 min read