arXivAI Research
Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark
重新審視 METR 時間跨度指標:以統計模型量化 AI 的真實任務能力
Re-analyzing METR benchmark data via splines and item-response theory reveals that AI task difficulty is non-linear, making a jump from 3 to 30 minutes far easier than from 30 minutes to 5 hours.
2 min read