Aivora
AnthropicAI SafetyIntermediate

Anthropic Expands Cyber Verification Program to Give Defenders the AI Advantage

Anthropic 擴大「網路安全驗證計畫」:放寬安全防護,為資安防守者提供強大 AI 武器

2 min read
Anthropic Expands Cyber Verification Program to Give Defenders the AI Advantage
The 30-second version

Because cybersecurity tools are dual-use, standard Claude models feature conservative safeguards that block most security-related tasks. To empower defenders, Anthropic merged its experimental initiatives into an expanded three-tier Cyber Verification Program (CVP): Defense, Red Team, and Specialized Access. This program provides vetted teams with reduced safeguards on flagship models like Claude Opus 5.5. Efficacy testing on CyScenarioBench showed that under the Red Team tier, Claude Opus 5.5 bypassed all standard safety blocks, achieving a 68% success rate in executing multi-stage cyber operations.

Key points

01

Three-Tier Access Structure

Introduces Defense, Red Team, and Specialized tiers, scaling safeguard reduction based on the organization's role and verification level.

02

Access to Premier Models

Qualified organizations gain access to elite models including Claude Opus 5.5, Claude Sonnet 5.5, and Claude Mythos 5.1.

03

Accelerated Vulnerability Finding

Partners using early versions of the program detected over 129,000 verified software vulnerabilities, accelerating finding speeds by months or years.

04

Privacy Safeguards Coming

Data retention is temporarily required to monitor misuse, but the upcoming Enterprise Frontier Safeguards (EFS) will enable zero-data retention.

How it works

CVP Access Tier Comparison
一般大眾模型防禦存取層級紅隊存取層級特許存取層級
Target大眾與一般企業企業防守團隊、開源維護者、個人研究員企業/政府紅隊、資安滲透測試廠商測試影響生命或市場之關鍵安全系統的特定組織
Primary Use一般任務、程式碼修補與檢視資安事件響應、逆向工程、漏洞分析經授權的滲透測試與對抗性安全演練航空系統、電網、政府網路等關鍵系統測試
Test Block Rate100%(首發對話即被安全防護阻擋)高度阻擋(92% 的測試在某階段被阻擋)0% 阻擋(對抗任務成功率約 68%)0% 阻擋(防護最少,需與政府深度審核)

Why it matters

AI safety guardrails are a double-edged sword, often blocking legitimate security professionals from using AI to patch systems. By creating a structured, verified channel, Anthropic safely delivers Claude's reasoning capabilities directly to defenders. This helps critical infrastructure, open-source maintainers, and enterprise teams automate tasks like reverse-engineering and penetration testing, shifting the strategic advantage back to systemic defense.

Who it affects

  • AI Developer
  • Enterprise Leader
  • Policy Maker
  • AI Researcher

How to use it

  1. 1Security Operations Center (SOC) tasks, incident response, reverse-engineering malware, and triaging security alerts.
  2. 2Authorized red-teaming and penetration testing against owned or authorized IT systems.
  3. 3Deep security testing of life-critical systems such as power grids, telecom networks, and flight control systems.

Limitations & caveats

  • Data retention is currently mandatory for monitoring, delaying zero-data retention benefits until EFS is launched later this fall.
  • The Red Team Access tier requires weeks of vetting and is currently restricted to verified organizations, excluding individual researchers.

Related

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
NVIDIA DeveloperAI Safety

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring

NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構

NVIDIA introduced the Open Agent Safety Platform, combining open-source OpenShell sandboxing and BlueField hardware DPUs to deliver independent, zero-trust monitoring and real-time intervention for autonomous agents.

2 min read
NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents
NVIDIA DeveloperAI Safety

NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents

NVIDIA OpenShell:不重寫程式碼,為 AI Agent 部署執行階段安全隔離防護

NVIDIA released OpenShell 0.1.0, an open-source runtime that secures AI agents using external sandboxing, supervisor monitoring, and formal policy analysis without rewriting agent code.

2 min read