Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
以「後設技能」為核心的 AI4AI:為 Agent 設計最佳執行環境的測試期學習架構
While traditional approaches focus on upgrading an agent's reasoning, this paper explores environment design. Under frozen weights, a Builder model constructs an execution harness for a Target agent. By introducing Meta-Skills—rules on when to support and what resources to provide—the Builder learns from Target feedback. Evaluated on Harness-Bench and NewtonBench, this approach outperforms no-skill setups by 8.95 percentage points and direct-delivery baselines by 12.02 percentage points, showing a path to system-level self-improvement.
Key points
AI-for-AI Harnessing
Explores how a Builder model can dynamically construct optimal execution environments (harnesses) for a Target Agent under frozen weights.
Innovative Meta-Skills
Translates experiences into principles specifying when support is needed and what resources to provide, making them reusable for unseen tasks.
Significant Gains
Improves macro-average performance by 8.95 percentage points over no-skill setups, and 12.02 points over direct delivery to the Target.
Weight-Free Improvement
Gains when the same model serves both roles suggest a path to system-level self-improvement through learning to build better environments.
How it works
Why it matters
Traditional AI self-improvement relies heavily on resource-intensive fine-tuning. This study demonstrates that optimizing the execution environment (harness) without modifying model weights can yield massive performance gains. It introduces a safer, more flexible, and cost-effective approach for compound AI systems, where models learn to act as capable "builders" to set up tailored environments for other agents.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Test-time environment optimization
- 2Weight-free self-improvement systems
Limitations & caveats
- Highly dependent on feedback from development sets; performance may degrade if dev and test tasks differ significantly.
- Introducing a Builder model adds dual-phase inference overhead and increases API invocation costs.
Related
Thinking Before Thinking: Scaling AI Agents with Meta-Reasoning
後設推理:讓 AI 代理在動手前先「思考如何思考」
This paper introduces agentic meta-reasoning, a framework using a controller to dynamically allocate compute budget, improving long-horizon task execution.
TokenCast: Accurate Token Consumption Forecasting for LLM Agents
TokenCast:精準預測 LLM Agent 執行過程中的 Token 消耗量
This paper introduces TokenCast, a real-time method that accurately forecasts dynamic LLM Agent token consumption within 32.8 ms without extra LLM calls, using composable cost representations.
Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)
以敘事狀態追蹤突破長篇小說生成:免微調的 NstAgent 框架
This study introduces NstAgent, a training-free agentic framework that dynamically tracks characters, past events, and future requirements to scale coherent LLM story generation up to 100K words.