Aivora
arXivAI AgentIntermediate

Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?

打造低成本 AI 工件:LLM Agent 是否具備「能力封裝」的本領?

2 min read
Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?
The 30-second version

Querying LLMs for millions of workload instances is prohibitively expensive. The authors propose "bottling"—where LLM agents autonomously build cheaper task-specific solutions (e.g., training a small model or writing a reusable script). They introduce BOTTLED, a benchmark to test agents under strict time, compute, and API budgets. Across 10 models, they found that zero-shot mastery does not guarantee bottling capability, as 48 of 60 runs fell below their zero-shot baseline. However, top-tier models like Opus 5 proved bottling's promise, cutting costs by ~657x while retaining 82% of its zero-shot F1 score.

Key points

01

The "Bottling" Concept

The ability of agents to autonomously convert general reasoning into task-specific, low-cost reusable solutions.

02

The Zero-Shot Performance Gap

Strong zero-shot scores do not translate to strong bottling; 48 of 60 runs fell below the models' zero-shot performance.

03

Massive Cost-Savings Potential

When successful, bottling yields huge savings; Opus 5 retained 82% of its F1 score at 657 times lower cost.

04

Struggling Against Baselines

Over half (31 of 60) of the runs failed to beat simpler, budget-equivalent small-model distillation baselines.

How it works

Zero-Shot Inference vs. Bottling Capabilities
直接調用 (Zero-Shot)能力封裝 (Bottling)
Core Strategy逐次向大型 LLM 進行昂貴查詢Agent 自動產出小模型或重用程式
Inference Cost高昂,隨查詢次數呈線性增長極低(如 Opus 5 可降低達 657 倍)
Task Quality100% 原始大型模型效能約 82% 效能(在保留品質下大幅降低成本)
Budget Limits無固定限制,預算隨規模耗盡受限於 BOTTLED 固定時間、算力與 Token

Why it matters

This research shifts the focus toward autonomous cost optimization for enterprise-scale LLM deployments. Traditionally, distilling models or creating specialized pipelines required heavy manual engineering. BOTTLED evaluates how well AI agents can automate this optimization process. Unlocking this capability could democratize industrial AI, making high-volume reasoning workloads incredibly cheap and accessible.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1High-volume classification: Allowing agents to automatically train specialized small classifiers, cutting massive API query fees.
  2. 2Automated workflow optimization: Internal enterprise agents writing code or fine-tuning lightweight models within strict resource constraints.

Limitations & caveats

  • Current agents lack consistency in bottling, with more than half failing to outperform standard distillation baselines.
  • The evaluation relies on static compute, token, and time budgets; dynamic real-world environments present trickier engineering challenges.

Related

AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model
arXivAI Agent

AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model

利用 Web 世界模型進行對抗訓練:AdvSim2Real 提升 Web Agent 抵禦適應性提示詞注入之能力

AdvSim2Real co-evolves a task curriculum, an injection adversary, and a web agent inside a frozen world model, boosting the robustness of a 4B agent against adaptive prompt injections by 33.6%.

2 min read
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
arXivAI Agent

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

點陣圖秒變任意比例流程圖!「One Figure, Every Canvas」以 Agent 協同管線自動排版且支援 draw.io 編輯

This research introduces an agentic pipeline that automatically reformats raster flowcharts into various aspect ratios while maintaining structural fidelity, outputting editable draw.io XML files.

2 min read