RECAST: Active Evidence Construction via Adaptive Routing and Computation
RECAST:透過自適應證據路由,為 LLM 計算出正確的上下文
Traditional RAG struggles when answers require aggregating or computing data across heterogeneous sources. RECAST solves this by modeling evidence construction as a sequential decision process. A lightweight RouterLM iteratively selects primitive operations or tasks a frozen CompilerLM to generate executable code. Once sufficient evidence is derived and accepted, a frozen AnswerLM outputs the final answer. Trained via SFT and GRPO, RECAST outperforms strong baselines by up to 15.9% and demonstrates exceptional zero-shot generalization across held-out tasks.
Key points
Active Evidence Derivation
Moves beyond passive retrieval to actively construct and synthesize evidence via iterative filtering, aggregation, and multi-step computation.
Router-Compiler-Answer Trio
Leverages a lightweight RouterLM for decisions, a frozen CompilerLM for code execution, and a frozen AnswerLM for final generation.
Reinforcement Learning via GRPO
Trained using SFT followed by Group Relative Policy Optimization (GRPO), enabling a 9B parameter RouterLM to outperform larger closed models.
Zero-Shot Generalization
Achieved a 15.0% average improvement over the strongest baseline on held-out benchmarks, demonstrating robust zero-shot task adaptation.
How it works
Why it matters
RECAST shifts the RAG paradigm from passive retrieval to active computation, which is essential for tasks requiring numerical reasoning and cross-source data aggregation. By proving that a 9B model optimized with GRPO can outperform a larger Gemini model, this research offers a cost-effective and highly accurate architectural blueprint for deploying specialized enterprise agents.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Financial analysis and multi-source database querying
- 2Complex Q&A requiring cross-document logic and numerical operations
- 3Automated data cleaning and multi-step code-generation reasoning
Limitations & caveats
- Highly dependent on secure code execution environments; syntax or logical errors in compiled code can disrupt the reasoning loop.
- The iterative routing and code compilation process introduces higher latency and computational overhead compared to single-shot retrieval.
Related

Microsoft Releases Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Real-Harness Training
微軟開源 Agent Lightning v1.0:僅 3,500 行程式碼,直接用生產環境 Harness 訓練 Agent 的強化學習框架
Microsoft Research Asia has introduced Agent Lightning v1.0, a lightweight open-source framework that trains AI agents using their actual deployment harnesses via a transparent LLM proxy.
Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?
打造低成本 AI 工件:LLM Agent 是否具備「能力封裝」的本領?
The study introduces the BOTTLED benchmark to evaluate if LLM agents can autonomously package their capabilities into low-cost, task-specific artifacts, revealing that strong zero-shot performance does not guarantee successful bottling.
AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model
利用 Web 世界模型進行對抗訓練:AdvSim2Real 提升 Web Agent 抵禦適應性提示詞注入之能力
AdvSim2Real co-evolves a task curriculum, an injection adversary, and a web agent inside a frozen world model, boosting the robustness of a 4B agent against adaptive prompt injections by 33.6%.