EvoDuet: Co-Evolving Web Search and Task Solving for LLM-Based Scientific Discovery
EvoDuet:搜尋與解題協同演化的雙層優化框架,突破 LLM 科學發現的知識瓶頸
Evolutionary search with LLMs often stalls when missing critical external knowledge. Simply adding search tools can lead to redundant page retrievals. EvoDuet addresses this by co-evolving search queries and task solutions under fixed parameters. In each iteration, a retrieval gate assesses the model's knowledge gap, an inner loop refines queries to rank documents, and an outer loop parallel-generates solution candidates. EvoDuet raised Gemini-3.8-Flash's OpenEvolve discovery gain from 61.3% to 82.3%, though smaller models like Qwen3.5-9B showed no improvement.
Key points
Bi-level Co-evolution
The inner loop refines search queries to retrieve and rank key documents, while the outer loop uses them to generate and evaluate solution candidates.
Active Retrieval Gate
A retrieval gate lets the LLM assess its own knowledge gap to decide whether to fetch new pages, reuse stored documents, or proceed without retrieval.
Significant Performance Gain
Across 21 tasks, it boosted Gemini-3.8-Flash's discovery gain to 82.3%, surpassing previous best scores on several complex algorithmic tasks.
Model Capability Threshold
The framework's success depends on the underlying LLM; smaller models like Qwen3.5-9B do not benefit from this co-evolutionary setup.
How it works
Why it matters
Complex scientific tasks require cutting-edge external knowledge. Traditional LLM agents often search inefficiently, fetching redundant data. EvoDuet proves that searching for knowledge and solving tasks can be jointly optimized without model fine-tuning. This offers a highly adaptive and efficient system architecture for AI-driven scientific discovery (AI for Science).
Who it affects
- AI Researcher
- AI Developer
- Student & Learner
How to use it
- 1Automated scientific algorithm and code optimization (e.g., Swap Reduction tasks)
- 2Complex agent task navigation requiring dynamic integration of up-to-date technical documentation
Limitations & caveats
- Highly dependent on the base LLM's reasoning and self-assessment; smaller or less capable models fail to yield performance gains.
- Retrieval efficacy is constrained by search engine quality; noisy initial search results may degrade inner-loop optimization efficiency.
Related

Overcoming Generative Recommender Latency: Deploying HSTU Models with NVIDIA Dynamo-Triton and PyTorch AOTI
突破生成式推薦延遲瓶頸:NVIDIA Dynamo-Triton 與 PyTorch AOTI 部署 HSTU 模型實戰
Learn how to deploy HSTU generative recommenders using NVIDIA Dynamo-Triton, PyTorch AOTI, and FlexKV caching to achieve up to a 5.93x speedup on Blackwell GPUs.
Ranking-PE: Prompt Optimization for Multimodal Clinical Diagnosis under Extreme Class Imbalance
臨床診斷 MLLM 提示詞優化:Ranking-PE 解決醫療資料極端不平衡問題
This paper introduces Ranking-PE, a ranking-aware prompt optimization framework that shifts MLLM adaptation from accuracy-based to AUROC-based ranking, resolving class imbalance in clinical diagnostics.
Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding
排除「時間捷徑」漏洞:非侵入式腦機介面解碼技術的新突破
Researchers revealed that recent breakthroughs in non-invasive brain-to-text decoding relied on a "timing shortcut" of word durations rather than actual brain signals. Their SimpleB2T method eliminates this shortcut, slashing the word error rate to 36.6%.