NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents
NVIDIA OpenShell:不重寫程式碼,為 AI Agent 部署執行階段安全隔離防護

While autonomous agents offer massive potential, giving them broad systems access introduces severe risks. NVIDIA OpenShell 0.1.0 mitigates this by placing runtime controls completely outside the agent's workload. Comprising a Gateway, Supervisor, and Sandbox, it provides kernel-level file isolation and deep-packet API inspection. OpenShell keeps actual credentials outside the workspace and uses formal policy logic to prevent agents from manipulating humans or AI reviewers into granting unauthorized privileges.
Key points
Out-of-workload enforcement
Enforcement occurs entirely outside the running agent workload, allowing quick integration without modifying existing agent code.
Three-tier architecture
Combines a central Gateway, distributed Supervisors, and kernel-isolated Sandboxes to sever direct, unauthorized network paths.
Granular API filtering
Goes beyond port blocking to inspect HTTP, GraphQL, and MCP traffic, enforcing specific rules like allow-reads and block-writes.
Formal policy verification
Uses a policy prover built on formal logic to ensure permissions stay bounded, preventing agents from socially engineering human reviewers.
How it works
Why it matters
As enterprises enter the era of autonomous AI agents executing code and API calls over long horizons, safety becomes paramount. OpenShell delivers a practical, open-source defense-in-depth framework. By shielding credentials and verifying runtime policies mathematically, it enables highly regulated industries—such as semiconductor design, robotics, and workflow automation—to safely unlock the potential of agent fleets.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Product Manager
How to use it
- 1Autonomous chip design: Cadence secures its ChipStack RTL design engineer agents using OpenShell's boundary controls.
- 2Physical robotics safety: Gecko Robotics implements OpenShell to govern critical, real-world actions made by edge-deployed agents.
- 3Enterprise automation: Slack is building an on-demand agent platform with OpenShell to safely execute internal automations.
Limitations & caveats
- Filesystem and process restrictions are bound at sandbox startup and cannot be modified dynamically during runtime.
- Joint policy analysis across multiple interacting agents is still an early-stage feature and undergoing active research.
Related
Compression Footprints as Security Signals: Defending Federated Learning Against Model Poisoning
壓縮足跡化身安全訊號:利用破壞性壓縮抵禦聯邦學習中的模型投毒攻擊
This study introduces CRAFT, a robust aggregation method that repurposes lossy compression distortions in Federated Learning into diagnostic footprints to detect and mitigate model-poisoning attacks.

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構
NVIDIA introduced the Open Agent Safety Platform, combining open-source OpenShell sandboxing and BlueField hardware DPUs to deliver independent, zero-trust monitoring and real-time intervention for autonomous agents.
Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users
以「信念自我蒸餾」提取使用者模型:揭示大語言模型對用戶意圖的內在表徵
Researchers introduce Belief Self-Distillation (BSD), a framework to read and write an LLM's implicit beliefs about its users, revealing how user-intent inference drives safety decisions and showing shared representation geometries across models.