arXivAI Coding
The Illusion of Compile Rate: Why Common Metrics Fail LLM-Based Vulnerability Repair
自動漏洞修復的指標迷思:為何「編譯率」與 CodeBLEU 無法真實反映 LLM 的修復能力
A study reveals that 'compile rate' and CodeBLEU fail as metrics for LLM-based vulnerability repair, often rewarding non-repairs, and proposes diff_F1 as a reliable, change-aware filtering alternative.
2 min read