Abstract
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, P., Suryanto, N., Zhang, S., & Naseer, M. (2026). KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards. https://omanscience.com/en/articles/kalibench-a-fine-grained-benchmark-for-cybersecurity-tool-use-on-kali-linux-with-runtime-free-verifiable-rewards
MLA 9
Li, Pengfei, et al. "KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards." https://omanscience.com/en/articles/kalibench-a-fine-grained-benchmark-for-cybersecurity-tool-use-on-kali-linux-with-runtime-free-verifiable-rewards.
Chicago (author–date)
Li, Pengfei, Naufal Suryanto, Sicheng Zhang, and Muzammal Naseer. 2026. "KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards." https://omanscience.com/en/articles/kalibench-a-fine-grained-benchmark-for-cybersecurity-tool-use-on-kali-linux-with-runtime-free-verifiable-rewards.
Harvard
Li, P., Suryanto, N., Zhang, S. and Naseer, M. (2026) 'KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards', Available at: https://omanscience.com/en/articles/kalibench-a-fine-grained-benchmark-for-cybersecurity-tool-use-on-kali-linux-with-runtime-free-verifiable-rewards.
Vancouver
Li P, Suryanto N, Zhang S, Naseer M. KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards. https://omanscience.com/en/articles/kalibench-a-fine-grained-benchmark-for-cybersecurity-tool-use-on-kali-linux-with-runtime-free-verifiable-rewards
IEEE
P. Li, N. Suryanto, S. Zhang, and M. Naseer, "KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards," https://omanscience.com/en/articles/kalibench-a-fine-grained-benchmark-for-cybersecurity-tool-use-on-kali-linux-with-runtime-free-verifiable-rewards.