Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
程式碼代理人的精簡文件有用嗎?研究揭示驚人否定結果:有了源碼,文件反而成多餘
The researchers introduced a "roundtrip benchmark" that evaluates code descriptions by seeing if code regenerated from them passes original tests, using this to optimize a high-fidelity documentation generator. However, testing across 10 repositories and two model families revealed a surprising negative result: when source code is present, neither static compact documentation nor retrieved context improves an agent's ability to resolve real issues beyond simply using the issue description itself.
Key points
Roundtrip Benchmark
Scores code descriptions based on whether code regenerated from them passes original tests, proving completeness, not length, drives fidelity.
Documentation Optimizer
Used the benchmark as an optimization signal to discover a description-writing prompt that achieves full fidelity and generalizes to unseen files.
Unexpected Negative Result
Tests across 2 model families and 10 repositories show that compact documentation does not improve an agent's real-world issue resolution.
Bound of Doc Utility
Defines the boundaries of when documentation actually helps, suggesting that documentation is redundant when full source code is accessible to agents.
How it works
Why it matters
It challenges the prevailing industry assumption that providing compact, high-density documentation to AI agents improves their coding performance. By proving this does not transfer to real-world issue resolution when source code is available, it saves researchers and developers from futile optimization efforts. This urges a paradigm shift in how we structure context and retrieval for LLM-based coding agents, refocusing on boundaries where documentation is actually necessary.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Evaluating the fidelity of existing codebase documentation by testing if regenerated code can pass original tests.
- 2Optimizing system prompts for internal code documentation generators to ensure completeness over briefness.
Limitations & caveats
- The negative transfer result is specific to scenarios where source code is present; documentation value may differ under closed-source or API-only constraints.
- We tested on two model families and ten repositories, meaning the boundary of when documentation helps could vary on larger or highly proprietary systems.
Related
CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents
CliffCompaction:長任務程式碼 Agent 的高效能 context 自動壓縮技術
CliffCompaction is an autocompaction technique for long-horizon coding agents that reduces token costs by up to 50% and prevents context drift by strictly truncating rather than rewriting original context.
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試
SWE-Serve is a new benchmark featuring 53 real-world SGLang tasks, designed to evaluate AI agents' capability to implement complex features and achieve production correctness across the inference serving stack.
The Illusion of Compile Rate: Why Common Metrics Fail LLM-Based Vulnerability Repair
自動漏洞修復的指標迷思:為何「編譯率」與 CodeBLEU 無法真實反映 LLM 的修復能力
A study reveals that 'compile rate' and CodeBLEU fail as metrics for LLM-based vulnerability repair, often rewarding non-repairs, and proposes diff_F1 as a reliable, change-aware filtering alternative.