Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?
打造低成本 AI 工件:LLM Agent 是否具備「能力封裝」的本領?
Querying LLMs for millions of workload instances is prohibitively expensive. The authors propose "bottling"—where LLM agents autonomously build cheaper task-specific solutions (e.g., training a small model or writing a reusable script). They introduce BOTTLED, a benchmark to test agents under strict time, compute, and API budgets. Across 10 models, they found that zero-shot mastery does not guarantee bottling capability, as 48 of 60 runs fell below their zero-shot baseline. However, top-tier models like Opus 5 proved bottling's promise, cutting costs by ~657x while retaining 82% of its zero-shot F1 score.
Key points
The "Bottling" Concept
The ability of agents to autonomously convert general reasoning into task-specific, low-cost reusable solutions.
The Zero-Shot Performance Gap
Strong zero-shot scores do not translate to strong bottling; 48 of 60 runs fell below the models' zero-shot performance.
Massive Cost-Savings Potential
When successful, bottling yields huge savings; Opus 5 retained 82% of its F1 score at 657 times lower cost.
Struggling Against Baselines
Over half (31 of 60) of the runs failed to beat simpler, budget-equivalent small-model distillation baselines.
How it works
| 直接調用 (Zero-Shot) | 能力封裝 (Bottling) | |
|---|---|---|
| Core Strategy | 逐次向大型 LLM 進行昂貴查詢 | Agent 自動產出小模型或重用程式 |
| Inference Cost | 高昂,隨查詢次數呈線性增長 | 極低(如 Opus 5 可降低達 657 倍) |
| Task Quality | 100% 原始大型模型效能 | 約 82% 效能(在保留品質下大幅降低成本) |
| Budget Limits | 無固定限制,預算隨規模耗盡 | 受限於 BOTTLED 固定時間、算力與 Token |
Why it matters
This research shifts the focus toward autonomous cost optimization for enterprise-scale LLM deployments. Traditionally, distilling models or creating specialized pipelines required heavy manual engineering. BOTTLED evaluates how well AI agents can automate this optimization process. Unlocking this capability could democratize industrial AI, making high-volume reasoning workloads incredibly cheap and accessible.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Product Manager
How to use it
- 1High-volume classification: Allowing agents to automatically train specialized small classifiers, cutting massive API query fees.
- 2Automated workflow optimization: Internal enterprise agents writing code or fine-tuning lightweight models within strict resource constraints.
Limitations & caveats
- Current agents lack consistency in bottling, with more than half failing to outperform standard distillation baselines.
- The evaluation relies on static compute, token, and time budgets; dynamic real-world environments present trickier engineering challenges.
Related
AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model
利用 Web 世界模型進行對抗訓練:AdvSim2Real 提升 Web Agent 抵禦適應性提示詞注入之能力
AdvSim2Real co-evolves a task curriculum, an injection adversary, and a web agent inside a frozen world model, boosting the robustness of a 4B agent against adaptive prompt injections by 33.6%.
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
點陣圖秒變任意比例流程圖!「One Figure, Every Canvas」以 Agent 協同管線自動排版且支援 draw.io 編輯
This research introduces an agentic pipeline that automatically reformats raster flowcharts into various aspect ratios while maintaining structural fidelity, outputting editable draw.io XML files.
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
MemPilot:為多模態 AI Agent 打造的按需動態記憶整理框架
MemPilot is a runtime memory curation framework for LLM agents that uses a reinforcement learning policy to dynamically balance performance, cost, and latency.