Aivora
arXivAI CodingIntermediate

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

程式碼代理人的精簡文件有用嗎?研究揭示驚人否定結果:有了源碼,文件反而成多餘

2 min read
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
The 30-second version

The researchers introduced a "roundtrip benchmark" that evaluates code descriptions by seeing if code regenerated from them passes original tests, using this to optimize a high-fidelity documentation generator. However, testing across 10 repositories and two model families revealed a surprising negative result: when source code is present, neither static compact documentation nor retrieved context improves an agent's ability to resolve real issues beyond simply using the issue description itself.

Key points

01

Roundtrip Benchmark

Scores code descriptions based on whether code regenerated from them passes original tests, proving completeness, not length, drives fidelity.

02

Documentation Optimizer

Used the benchmark as an optimization signal to discover a description-writing prompt that achieves full fidelity and generalizes to unseen files.

03

Unexpected Negative Result

Tests across 2 model families and 10 repositories show that compact documentation does not improve an agent's real-world issue resolution.

04

Bound of Doc Utility

Defines the boundaries of when documentation actually helps, suggesting that documentation is redundant when full source code is accessible to agents.

How it works

Roundtrip Benchmark and Doc Optimization Flow
Opt SignalSource CodeOptimized GenCompact DocCode RegenRun TestsFidelity Score

Why it matters

It challenges the prevailing industry assumption that providing compact, high-density documentation to AI agents improves their coding performance. By proving this does not transfer to real-world issue resolution when source code is available, it saves researchers and developers from futile optimization efforts. This urges a paradigm shift in how we structure context and retrieval for LLM-based coding agents, refocusing on boundaries where documentation is actually necessary.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Evaluating the fidelity of existing codebase documentation by testing if regenerated code can pass original tests.
  2. 2Optimizing system prompts for internal code documentation generators to ensure completeness over briefness.

Limitations & caveats

  • The negative transfer result is specific to scenarios where source code is present; documentation value may differ under closed-source or API-only constraints.
  • We tested on two model families and ten repositories, meaning the boundary of when documentation helps could vary on larger or highly proprietary systems.

Related

CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents
arXivAI Coding

CliffCompaction: Cost-Efficient Context Compaction for Long-Horizon Coding Agents

CliffCompaction:長任務程式碼 Agent 的高效能 context 自動壓縮技術

CliffCompaction is an autocompaction technique for long-horizon coding agents that reduces token costs by up to 50% and prevents context drift by strictly truncating rather than rewriting original context.

2 min read
SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving
arXivAI Coding

SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving

SWE-Serve:首個針對「生產級推論服務」的 AI Agent 軟體工程基準測試

SWE-Serve is a new benchmark featuring 53 real-world SGLang tasks, designed to evaluate AI agents' capability to implement complex features and achieve production correctness across the inference serving stack.

2 min read
The Illusion of Compile Rate: Why Common Metrics Fail LLM-Based Vulnerability Repair
arXivAI Coding

The Illusion of Compile Rate: Why Common Metrics Fail LLM-Based Vulnerability Repair

自動漏洞修復的指標迷思:為何「編譯率」與 CodeBLEU 無法真實反映 LLM 的修復能力

A study reveals that 'compile rate' and CodeBLEU fail as metrics for LLM-based vulnerability repair, often rewarding non-repairs, and proposes diff_F1 as a reliable, change-aware filtering alternative.

2 min read