KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸
Deploying LLMs in security workflows requires precise translation of intent into strict CLI commands. KaliBench addresses this with 8,504 query-command pairs spanning 1,642 Kali Linux tools. Evaluations reveal that open-weight models struggle in this domain, with none exceeding 42% exact accuracy without explicit tool hints. To solve this, KaliBench introduces runtime-free verifiable rewards for training. Implementing this mechanism via SFT and reinforcement learning significantly boosts an 8B model to perform on par with a 685B MoE model.
Key points
Precise CLI Translation
Evaluates the LLM's ability to output exact terminal commands without minor syntax, flag, or ordering errors.
Large-Scale Cybersecurity Dataset
Features 8,504 query-command pairs across 1,642 tools, 23 capability dimensions, and 5 security phases.
Multi-Stage Verification
Combines LLM validation, sandboxed execution, and human review to ensure absolute real-world executability.
Runtime-Free Verifiable Rewards
Enables deterministic reward signals during training without spinning up active runtime environments.
How it works
| 一般開源模型 (Open-Weight) | KaliBench 訓練後 8B 模型 | 685B MoE 旗艦模型 | |
|---|---|---|---|
| Command Accuracy | 低於 42% (無提示) | 顯著提升並媲美 MoE | 基準測試最高水平 |
| Hardware & Compute Cost | 中等 | 極低 (8B 規模) | 極高 (685B 規模) |
| RL Optimization Pathway | 缺乏標準反饋訊號 | 免執行期可驗證獎勵 | 訓練極度昂貴 |
Why it matters
In practical SecOps or red-teaming, a single flag error can break automated scripts. KaliBench highlights this severe limitation in current models (<42% exact accuracy) and introduces a runtime-free verifiable reward mechanism. This allows organizations to train highly efficient, localized 8B models to match the precision of a 685B MoE flagship model, slashing deployment costs and accelerating safe, automated terminal operations.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Training autonomous SecOps and penetration testing agents
- 2Benchmarking domain-specific security LLMs on tool-use accuracy
- 3Low-overhead reinforcement learning for CLI translation tasks
Limitations & caveats
- Scoped specifically to the Kali Linux toolset, meaning transferability to other platforms remains untested.
- Focuses on single-step CLI generation and canonicalization, rather than stateful, multi-turn sequence planning.
Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料
AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架
VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.

Google Announces Gemini 4 Argon: Frontier Model with 1M Output Tokens and Long-Horizon Reasoning
Google 發表新一代前沿模型 Gemini 4 Argon:具備 100 萬 Token 輸出與強大自主 Agent 推理能力
Google has unveiled Gemini 4 Argon, featuring an industry-leading 1-million output token limit designed to sustain deep reasoning across complex, long-horizon workflows.