Aivora
arXivAI CodingIntermediate

FrugalEvo: Cost-Aware LLM-Guided Program Evolution via Dual-Model Collaboration

FrugalEvo:雙模型協作的成本感知型 LLM 程式最佳化演化框架

2 min read
FrugalEvo: Cost-Aware LLM-Guided Program Evolution via Dual-Model Collaboration
The 30-second version

Traditional LLM-guided program evolution focuses on performance over fixed iterations, ignoring high API costs. FrugalEvo introduces cost-awareness to this process using a dual-LLM approach: a stronger, higher-cost LLM (e.g., GPT-5.6 Terra) conceptualizes high-level strategies, while a cheaper LLM (e.g., Luna) implements and refines the code. By optimizing prompt design to maximize KV cache reuse, it drastically cuts API expenses. To evaluate efficiency under a budget, the authors propose the Budget-Aware Area Under the Curve (BA-AUC) metric. In tasks like circle packing, FrugalEvo achieves state-of-the-art results for just $0.55–$1.68, down from the ~$50 spent by multi-agent baselines.

Key points

01

Dual-LLM Strategy

Combines a powerful LLM for strategy exploration with a cheaper LLM for code implementation and refinement to balance depth and API budget.

02

Cache-Efficient Evolution

Optimizes harness and prompt structures to maximize prefix sharing across evolutionary steps, significantly boosting KV cache reuse.

03

Budget-Aware Metric (BA-AUC)

Introduces Budget-Aware Area Under the Curve (BA-AUC) to measure optimization quality relative to cumulative monetary spending rather than just iteration counts.

04

Massive Cost Savings

Achieves new SOTA on circle packing for just $0.55 to $1.68, compared to multi-agent baselines costing around $50 on average.

How it works

FrugalEvo Cost-Aware Dual-LLM Flow
Under BudgetStrong LLMFormulate StrategyCheap LLMWrite & Refine CodeKV Cache ReuseBudget Check

Why it matters

As automated coding agents rely heavily on iterative LLM calls, compounding API costs are a major barrier. FrugalEvo demonstrates that cost-aware algorithm design—leveraging heterogeneous LLMs and prompt-level cache engineering—can slash operational costs by over 90% while matching SOTA performance, making large-scale automated program optimization practical for real-world deployment.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1Algorithmic optimization: Automatically refining complex mathematical or systems algorithms (e.g., packing problems) under a strict API budget.
  2. 2Low-cost code synthesis: Separating high-level logic design from repetitive coding tasks to lower developers' testing costs.

Limitations & caveats

  • Relies heavily on LLM APIs with efficient prefix caching; benefits may decrease if the underlying provider lacks this feature.
  • Cheaper models may struggle to accurately implement highly complex logic in tasks requiring continuous advanced reasoning.

Related

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
arXivAI Coding

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

程式碼代理人的精簡文件有用嗎?研究揭示驚人否定結果:有了源碼,文件反而成多餘

This study investigates if compact natural-language documentation helps coding agents resolve software issues, finding that despite building high-fidelity documentation, it fails to improve real-world issue resolution over using raw source code and issue descriptions alone.

2 min read
CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents
arXivAI Coding

CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents

CliffCompaction:長任務程式碼 Agent 的高效能 context 自動壓縮技術

CliffCompaction is an autocompaction technique for long-horizon coding agents that reduces token costs by up to 50% and prevents context drift by strictly truncating rather than rewriting original context.

2 min read
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
arXivAI Coding

SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving

SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試

SWE-Serve is a new benchmark featuring 53 real-world SGLang tasks, designed to evaluate AI agents' capability to implement complex features and achieve production correctness across the inference serving stack.

2 min read