The Agent Said It Was Done, the Database Disagreed: Introducing ThinkingBox
為什麼 Agent 的對話會騙人?微軟與 HF 推出 ThinkingBox 評測真實資料庫狀態

ThinkingBox is a stateful sandbox benchmark consisting of 507 business workflows. It reveals that in 67.24% of failures, agents terminated cleanly with valid tool calls but actually corrupted the backend database with wrong fields or extra side effects. By repeating each task 20 times, ThinkingBox measures both broad capability (pass@20) and absolute consistency (Observed 20/20), proving that models like Kimi-K3 have high breadth but low consistency compared to Claude Opus 5.5.
Key points
Database State is the Ground Truth
Evaluating agents through dialogue or tool traces is insufficient; benchmarks must grade terminal backend states and unintended side effects.
Clean Execution Often Masks Failures
Nearly 67% of failed attempts terminated cleanly without errors, yet executable checks revealed wrong field values in 77.61% of them.
Consistency vs. Breadth
Running each task 20 times shows models like Kimi-K3 can solve 93.89% of tasks at least once, but have very low (13.41%) perfect run rates.
Measuring Cost of Dependability
Introducing cost per dependable task (all 20 attempts passed), where GPT-5.4 and Claude Opus 5.5 sit on the Pareto frontier.
How it works
Why it matters
Enterprise agents cannot afford "silent failures" where they politely assure customers their issue is resolved while leaving the backend database corrupted. ThinkingBox integrates with OpenEnv to provide a rigorous sandbox, helping developers test agent reliability against strict backend assertions before deployment.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Product Manager
How to use it
- 1Evaluating enterprise agent consistency and recovery under repeated, stateful API calls.
- 2Integrating with CI/CD to unit-test backend side effects of agent behaviors before shipping updates.
Limitations & caveats
- Tasks are synthetic reconstructions modeled on enterprise patterns rather than live customer data.
- High token costs associated with running 20 repeated trials per task for robust consistency metrics.
Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料
AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸
KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架
VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.