Aivora
Hugging FaceAI SafetyIntermediate

Solving the AI Evaluation Reproducibility Crisis: UK AISI and EvalEval Partner for Standardized Benchmarks

解決 AI 評測重現難題:UK AISI 與 EvalEval 聯手推動基準測試標準化

2 min read
Solving the AI Evaluation Reproducibility Crisis: UK AISI and EvalEval Partner for Standardized Benchmarks
The 30-second version

As AI deployment accelerates, benchmarks suffer from inconsistent reporting and high reproduction costs. To address this, UK AISI and EvalEval are implementing the "Every Eval Ever" (EEE) schema and "Evaluation Cards" platform. AISI has released verified evaluation data across five key benchmarks (including HealthBench and FrontierMath) for frontier models like Claude Opus 4 and GPT-5, demonstrating how setup choices and inference compute alter reported performance.

Key points

01

Standardizing Eval Reporting

Introduces the "Every Eval Ever" (EEE) shared schema to consolidate fragmented evaluation results and metadata into structured "Evaluation Cards".

02

High-Fidelity Transparency

Emphasizes transcript-level transparency, which is critical for analyzing, diagnosing model behavior, and ensuring reproducibility.

03

Impact of Inference Compute

AISI research shows that frontier model performance (e.g., on Humanity's Last Exam) heavily depends on inference-time compute and evaluation protocols.

04

Open Reference Datasets

Releases verified data for five benchmarks, covering six frontier models including the Claude Opus 4 and GPT-5 series as open reference points.

How it works

Evaluation Standardization and Reproducibility Flow
StandardizeIntegrateGenerateReproduce & AnalyzeRaw Eval DataConfigs & TranscriptsEEE SchemaEvaluation CardPolicy & Researchers

Why it matters

Current AI evaluations lack standardization, making it hard to discern if similar scores came from vastly different conditions. By using the EEE schema and Evaluation Cards, researchers, developers, and policymakers can compare findings without the prohibitive cost of re-running evaluations. This establishes a reliable, interoperable foundation for AI governance and scientific safety assessments.

Who it affects

  • AI Researcher
  • Policy Maker
  • AI Developer
  • Enterprise Leader

How to use it

  1. 1AI Meta-Research: Policy and governance researchers use Evaluation Cards to compare safety metrics across different platforms and understand how setup choices alter outcomes.
  2. 2Standardized Benchmark Reporting: Evaluation developers publish new benchmarks using the Every Eval Ever schema to ensure the community can fully replicate their experimental setups.

Limitations & caveats

  • Verified data is currently only fully published for five main benchmarks (e.g., HealthBench, FrontierMath) and two cybersecurity evaluations.
  • The ecosystem's success relies heavily on voluntary adoption of the EEE schema by model developers and evaluation organizations.

Related

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
NVIDIA DeveloperAI Safety

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring

NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構

NVIDIA introduced the Open Agent Safety Platform, combining open-source OpenShell sandboxing and BlueField hardware DPUs to deliver independent, zero-trust monitoring and real-time intervention for autonomous agents.

2 min read
NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents
NVIDIA DeveloperAI Safety

NVIDIA OpenShell: Enforcing Secure Runtime Controls and Sandboxing for AI Agents

NVIDIA OpenShell:不重寫程式碼,為 AI Agent 部署執行階段安全隔離防護

NVIDIA released OpenShell 0.1.0, an open-source runtime that secures AI agents using external sandboxing, supervisor monitoring, and formal policy analysis without rewriting agent code.

2 min read