Aivora

#item-response-theory

Item Response Theory

1 article

Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark
arXivAI Research

Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark

重新審視 METR 時間跨度指標:以統計模型量化 AI 的真實任務能力

Re-analyzing METR benchmark data via splines and item-response theory reveals that AI task difficulty is non-linear, making a jump from 3 to 30 minutes far easier than from 30 minutes to 5 hours.

2 min read