Aivora
NVIDIA DeveloperAI SafetyIntermediate

NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents

NVIDIA OpenShell:不重寫程式碼,為 AI Agent 部署執行階段安全隔離防護

2 min read
NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents
The 30-second version

While autonomous agents offer massive potential, giving them broad systems access introduces severe risks. NVIDIA OpenShell 0.1.0 mitigates this by placing runtime controls completely outside the agent's workload. Comprising a Gateway, Supervisor, and Sandbox, it provides kernel-level file isolation and deep-packet API inspection. OpenShell keeps actual credentials outside the workspace and uses formal policy logic to prevent agents from manipulating humans or AI reviewers into granting unauthorized privileges.

Key points

01

Out-of-workload enforcement

Enforcement occurs entirely outside the running agent workload, allowing quick integration without modifying existing agent code.

02

Three-tier architecture

Combines a central Gateway, distributed Supervisors, and kernel-isolated Sandboxes to sever direct, unauthorized network paths.

03

Granular API filtering

Goes beyond port blocking to inspect HTTP, GraphQL, and MCP traffic, enforcing specific rules like allow-reads and block-writes.

04

Formal policy verification

Uses a policy prover built on formal logic to ensure permissions stay bounded, preventing agents from socially engineering human reviewers.

How it works

NVIDIA OpenShell 3-Tier Security & Inspection Architecture
Deploy policies & lifecycleRuns isolated insideOutbound HTTP/MCP requestApprove & attach credentialsOpenShell GatewayAgent WorkloadOpenShell SupervisorOpenShell SandboxExternal Services

Why it matters

As enterprises enter the era of autonomous AI agents executing code and API calls over long horizons, safety becomes paramount. OpenShell delivers a practical, open-source defense-in-depth framework. By shielding credentials and verifying runtime policies mathematically, it enables highly regulated industries—such as semiconductor design, robotics, and workflow automation—to safely unlock the potential of agent fleets.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1Autonomous chip design: Cadence secures its ChipStack RTL design engineer agents using OpenShell's boundary controls.
  2. 2Physical robotics safety: Gecko Robotics implements OpenShell to govern critical, real-world actions made by edge-deployed agents.
  3. 3Enterprise automation: Slack is building an on-demand agent platform with OpenShell to safely execute internal automations.

Limitations & caveats

  • Filesystem and process restrictions are bound at sandbox startup and cannot be modified dynamically during runtime.
  • Joint policy analysis across multiple interacting agents is still an early-stage feature and undergoing active research.

Related

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
NVIDIA DeveloperAI Safety

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring

NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構

NVIDIA introduced the Open Agent Safety Platform, combining open-source OpenShell sandboxing and BlueField hardware DPUs to deliver independent, zero-trust monitoring and real-time intervention for autonomous agents.

2 min read
Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users
arXivAI Safety

Extracting User Models via Belief Self-Distillation: How LLMs Form and Use Beliefs About Users

以「信念自我蒸餾」提取使用者模型:揭示大語言模型對用戶意圖的內在表徵

Researchers introduce Belief Self-Distillation (BSD), a framework to read and write an LLM's implicit beliefs about its users, revealing how user-intent inference drives safety decisions and showing shared representation geometries across models.

2 min read