Thinking Before Thinking: Scaling AI Agents with Meta-Reasoning
後設推理:讓 AI 代理在動手前先「思考如何思考」
As AI agents face longer and more complex tasks, controlling execution paths becomes critical. This paper proposes 'agentic meta-reasoning', which separates execution into task-level 'workers' and a high-level 'controller'. Instead of replaying full histories, the controller carries a compact summary, evaluates options against the remaining budget, and dispatches work using persistent memory. Across long-horizon benchmarks, this approach continued to scale and improve performance where direct control methods typically plateaued.
Key points
Two-Tiered Architecture
Separates execution into task-level workers and a high-level controller responsible for planning and option dispatching.
Compact History Tracking
The controller carries a compact account of the run between decisions, relying on persistent memory for context rather than full history replays.
Scaling with Budget
Unlike direct control methods that plateau, meta-reasoning performance keeps scaling as the inference compute budget increases.
Superior Benchmarks
On ProgramBench, GPT-5.5 with meta-reasoning achieved 71.5% compared to 58.0% for Codex, showing notable gains in long-horizon tasks.
How it works
Why it matters
This research demonstrates that scaling inference compute is most effective when managed by a structured control layer. By actively evaluating budget, deciding when to reuse artifacts, and discarding dead ends, agents can handle long-horizon coding and scientific workflows much more reliably. It offers a new architectural paradigm for building robust, autonomous AI agents.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Autonomous software engineering and code reconstruction, with a controller orchestrating workers to debug and manage codebases.
- 2Long-horizon scientific exploration and mathematical proof generation, dynamically exploring proof paths under a set budget.
Limitations & caveats
- Small-budget overhead: The framework's coordination overhead can degrade performance when operating under very small compute budgets.
- Dependency on base models: The quality of meta-reasoning and execution remains constrained by the baseline capabilities of the underlying frontier models.
Related
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
以「後設技能」為核心的 AI4AI:為 Agent 設計最佳執行環境的測試期學習架構
This study introduces a test-time AI-for-AI framework where a Builder model learns Meta-Skills to design optimal execution environments (harnesses) for a target agent without training weights, boosting performance on unseen tasks.
TokenCast: Accurate Token Consumption Forecasting for LLM Agents
TokenCast:精準預測 LLM Agent 執行過程中的 Token 消耗量
This paper introduces TokenCast, a real-time method that accurately forecasts dynamic LLM Agent token consumption within 32.8 ms without extra LLM calls, using composable cost representations.
Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)
以敘事狀態追蹤突破長篇小說生成:免微調的 NstAgent 框架
This study introduces NstAgent, a training-free agentic framework that dynamically tracks characters, past events, and future requirements to scale coherent LLM story generation up to 100K words.