Anthropic Expands Cyber Verification Program to Give Defenders the AI Advantage
Anthropic 擴大「網路安全驗證計畫」:放寬安全防護,為資安防守者提供強大 AI 武器
Because cybersecurity tools are dual-use, standard Claude models feature conservative safeguards that block most security-related tasks. To empower defenders, Anthropic merged its experimental initiatives into an expanded three-tier Cyber Verification Program (CVP): Defense, Red Team, and Specialized Access. This program provides vetted teams with reduced safeguards on flagship models like Claude Opus 5.5. Efficacy testing on CyScenarioBench showed that under the Red Team tier, Claude Opus 5.5 bypassed all standard safety blocks, achieving a 68% success rate in executing multi-stage cyber operations.
Key points
Three-Tier Access Structure
Introduces Defense, Red Team, and Specialized tiers, scaling safeguard reduction based on the organization's role and verification level.
Access to Premier Models
Qualified organizations gain access to elite models including Claude Opus 5.5, Claude Sonnet 5.5, and Claude Mythos 5.1.
Accelerated Vulnerability Finding
Partners using early versions of the program detected over 129,000 verified software vulnerabilities, accelerating finding speeds by months or years.
Privacy Safeguards Coming
Data retention is temporarily required to monitor misuse, but the upcoming Enterprise Frontier Safeguards (EFS) will enable zero-data retention.
How it works
| 一般大眾模型 | 防禦存取層級 | 紅隊存取層級 | 特許存取層級 | |
|---|---|---|---|---|
| Target | 大眾與一般企業 | 企業防守團隊、開源維護者、個人研究員 | 企業/政府紅隊、資安滲透測試廠商 | 測試影響生命或市場之關鍵安全系統的特定組織 |
| Primary Use | 一般任務、程式碼修補與檢視 | 資安事件響應、逆向工程、漏洞分析 | 經授權的滲透測試與對抗性安全演練 | 航空系統、電網、政府網路等關鍵系統測試 |
| Test Block Rate | 100%(首發對話即被安全防護阻擋) | 高度阻擋(92% 的測試在某階段被阻擋) | 0% 阻擋(對抗任務成功率約 68%) | 0% 阻擋(防護最少,需與政府深度審核) |
Why it matters
AI safety guardrails are a double-edged sword, often blocking legitimate security professionals from using AI to patch systems. By creating a structured, verified channel, Anthropic safely delivers Claude's reasoning capabilities directly to defenders. This helps critical infrastructure, open-source maintainers, and enterprise teams automate tasks like reverse-engineering and penetration testing, shifting the strategic advantage back to systemic defense.
Who it affects
- AI Developer
- Enterprise Leader
- Policy Maker
- AI Researcher
How to use it
- 1Security Operations Center (SOC) tasks, incident response, reverse-engineering malware, and triaging security alerts.
- 2Authorized red-teaming and penetration testing against owned or authorized IT systems.
- 3Deep security testing of life-critical systems such as power grids, telecom networks, and flight control systems.
Limitations & caveats
- Data retention is currently mandatory for monitoring, delaying zero-data retention benefits until EFS is launched later this fall.
- The Red Team Access tier requires weeks of vetting and is currently restricted to verified organizations, excluding individual researchers.
Related
Compression Footprints as Security Signals: Defending Federated Learning Against Model Poisoning
壓縮足跡化身安全訊號:利用破壞性壓縮抵禦聯邦學習中的模型投毒攻擊
This study introduces CRAFT, a robust aggregation method that repurposes lossy compression distortions in Federated Learning into diagnostic footprints to detect and mitigate model-poisoning attacks.

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構
NVIDIA introduced the Open Agent Safety Platform, combining open-source OpenShell sandboxing and BlueField hardware DPUs to deliver independent, zero-trust monitoring and real-time intervention for autonomous agents.

NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents
NVIDIA OpenShell:不重寫程式碼,為 AI Agent 部署執行階段安全隔離防護
NVIDIA released OpenShell 0.1.0, an open-source runtime that secures AI agents using external sandboxing, supervisor monitoring, and formal policy analysis without rewriting agent code.