Aivora
arXivAI AgentIntermediate

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸

2 min read
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
The 30-second version

Deploying LLMs in security workflows requires precise translation of intent into strict CLI commands. KaliBench addresses this with 8,504 query-command pairs spanning 1,642 Kali Linux tools. Evaluations reveal that open-weight models struggle in this domain, with none exceeding 42% exact accuracy without explicit tool hints. To solve this, KaliBench introduces runtime-free verifiable rewards for training. Implementing this mechanism via SFT and reinforcement learning significantly boosts an 8B model to perform on par with a 685B MoE model.

Key points

01

Precise CLI Translation

Evaluates the LLM's ability to output exact terminal commands without minor syntax, flag, or ordering errors.

02

Large-Scale Cybersecurity Dataset

Features 8,504 query-command pairs across 1,642 tools, 23 capability dimensions, and 5 security phases.

03

Multi-Stage Verification

Combines LLM validation, sandboxed execution, and human review to ensure absolute real-world executability.

04

Runtime-Free Verifiable Rewards

Enables deterministic reward signals during training without spinning up active runtime environments.

How it works

Model Performance Comparison on KaliBench
一般開源模型 (Open-Weight)KaliBench 訓練後 8B 模型685B MoE 旗艦模型
Command Accuracy低於 42% (無提示)顯著提升並媲美 MoE基準測試最高水平
Hardware & Compute Cost中等極低 (8B 規模)極高 (685B 規模)
RL Optimization Pathway缺乏標準反饋訊號免執行期可驗證獎勵訓練極度昂貴

Why it matters

In practical SecOps or red-teaming, a single flag error can break automated scripts. KaliBench highlights this severe limitation in current models (<42% exact accuracy) and introduces a runtime-free verifiable reward mechanism. This allows organizations to train highly efficient, localized 8B models to match the precision of a 685B MoE flagship model, slashing deployment costs and accelerating safe, automated terminal operations.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Training autonomous SecOps and penetration testing agents
  2. 2Benchmarking domain-specific security LLMs on tool-use accuracy
  3. 3Low-overhead reinforcement learning for CLI translation tasks

Limitations & caveats

  • Scoped specifically to the Kali Linux toolset, meaning transferability to other platforms remains untested.
  • Focuses on single-step CLI generation and canonicalization, rather than stateful, multi-turn sequence planning.

Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
Hugging FaceAI Agent

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.

2 min read
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
arXivAI Agent

VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness

VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架

VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.

2 min read