FrugalEvo: Cost-Aware LLM-Guided Program Evolution via Dual-Model Collaboration
FrugalEvo:雙模型協作的成本感知型 LLM 程式最佳化演化框架
Traditional LLM-guided program evolution focuses on performance over fixed iterations, ignoring high API costs. FrugalEvo introduces cost-awareness to this process using a dual-LLM approach: a stronger, higher-cost LLM (e.g., GPT-5.6 Terra) conceptualizes high-level strategies, while a cheaper LLM (e.g., Luna) implements and refines the code. By optimizing prompt design to maximize KV cache reuse, it drastically cuts API expenses. To evaluate efficiency under a budget, the authors propose the Budget-Aware Area Under the Curve (BA-AUC) metric. In tasks like circle packing, FrugalEvo achieves state-of-the-art results for just $0.55–$1.68, down from the ~$50 spent by multi-agent baselines.
Key points
Dual-LLM Strategy
Combines a powerful LLM for strategy exploration with a cheaper LLM for code implementation and refinement to balance depth and API budget.
Cache-Efficient Evolution
Optimizes harness and prompt structures to maximize prefix sharing across evolutionary steps, significantly boosting KV cache reuse.
Budget-Aware Metric (BA-AUC)
Introduces Budget-Aware Area Under the Curve (BA-AUC) to measure optimization quality relative to cumulative monetary spending rather than just iteration counts.
Massive Cost Savings
Achieves new SOTA on circle packing for just $0.55 to $1.68, compared to multi-agent baselines costing around $50 on average.
How it works
Why it matters
As automated coding agents rely heavily on iterative LLM calls, compounding API costs are a major barrier. FrugalEvo demonstrates that cost-aware algorithm design—leveraging heterogeneous LLMs and prompt-level cache engineering—can slash operational costs by over 90% while matching SOTA performance, making large-scale automated program optimization practical for real-world deployment.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Product Manager
How to use it
- 1Algorithmic optimization: Automatically refining complex mathematical or systems algorithms (e.g., packing problems) under a strict API budget.
- 2Low-cost code synthesis: Separating high-level logic design from repetitive coding tasks to lower developers' testing costs.
Limitations & caveats
- Relies heavily on LLM APIs with efficient prefix caching; benefits may decrease if the underlying provider lacks this feature.
- Cheaper models may struggle to accurately implement highly complex logic in tasks requiring continuous advanced reasoning.
Related
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
程式碼代理人的精簡文件有用嗎?研究揭示驚人否定結果:有了源碼,文件反而成多餘
This study investigates if compact natural-language documentation helps coding agents resolve software issues, finding that despite building high-fidelity documentation, it fails to improve real-world issue resolution over using raw source code and issue descriptions alone.
CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents
CliffCompaction:長任務程式碼 Agent 的高效能 context 自動壓縮技術
CliffCompaction is an autocompaction technique for long-horizon coding agents that reduces token costs by up to 50% and prevents context drift by strictly truncating rather than rewriting original context.
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試
SWE-Serve is a new benchmark featuring 53 real-world SGLang tasks, designed to evaluate AI agents' capability to implement complex features and achieve production correctness across the inference serving stack.