NeurIPS 2026Evaluations and Datasets Track

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Pengfei Li1,∗Naufal Suryanto1,∗Sicheng Zhang1Muzammal Naseer1,2

1 Khalifa University2 University of Western Australia ∗ Equal contribution

Paper: arXiv version coming soon.

8,504
Verified query–command pairs
1,642
Kali Linux tools
23
Capability dimensions
5
Security phases

Benchmark overview

KaliBench separates tool selection from argument construction to reveal where command generation fails.

KaliBench evaluates how well language models translate natural-language cybersecurity requests into executable Kali Linux commands. It isolates tool invocation from end-to-end agent performance and measures errors at the level of individual command components.

The dataset is grounded in official tool documentation, checked through sandboxed execution and human review, and split into 3,504 training pairs and 5,000 evaluation pairs. The same deterministic scoring provides verifiable rewards for training without executing model-generated commands.

Read the abstract

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts’ intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs’ ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag–value bindings, or argument misordering can invalidate execution.

We introduce KaliBench, a fine-grained benchmark and dataset for natural-language–to–CLI translation on Kali Linux, comprising 8,504 query–command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement.

Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

Explore benchmark examples

23 dimensions · 5 security phases

Natural-language request

Scan host 192.168.1.100 using TCP Connect scan on ports 21 through 25 and 80, and display only open ports.

Released dataset example
Test split · sample_8052

Ground-truth command

nmap --open -sT -p 21-25,80 192.168.1.100

Tool

nmap

Positional arguments

  1. 192.168.1.100

Optional arguments

Key / aliases Value
--open nullFlag without a value
-sT nullFlag without a value
-p 21-25,80

Representative dataset examples. Commands are displayed, never run.

Download examples (JSON)

Key findings

24 open-weight model configurations, evaluated on 5,000 queries.

41.3%

Exact commands remain difficult

The strongest evaluated open-weight model gets fewer than half of the commands exactly right without tool hints.

22.3→73.1%

Tool documentation improves accuracy

Average exact-command accuracy rises sharply with tool documentation. Argument construction is the main bottleneck.

+7.5points

Targeted training improves an 8B model

SFT followed by GRPO raises average total score from 71.7 to 79.2, within 1.0 point of DeepSeek-V3.2.

Source: Table 1. The 41.3% result applies to the evaluated open-weight models; additional systems are reported separately below.

Results across three evaluation settings

Select how much tool information the model receives, then compare performance by metric.

Unrestricted results

All 24 models

Scroll horizontally to compare all metrics

Open-weight model results from Table 1 of the paper. All scores are percentages. Sorted by the selected metric.
Model Exact correct ↓ Size (B) Tool acc. Optional F1 Positional F1 Total score
GLM-5.2† 41.3 753 86.3 75.3 51.3 71.0
DeepSeek-V3.2† 33.0 685 / 37 75.8 57.0 68.8 67.2
RedSage-K (SFT+GRPO)†Ours 32.2 8 77.9 56.1 76.9 70.3
RedSage-K (SFT)Ours 30.6 8 78.1 52.8 66.9 65.9
RedSage-K (GRPO)†Ours 28.1 8 76.5 51.7 73.9 67.4
Qwen3-Coder-Next† 26.2 80 / 3 74.1 48.8 63.9 62.3
Qwen2.5-Ins 25.4 72 76.0 47.7 67.0 63.6
Qwen3† 24.0 32 72.4 47.2 65.3 61.6
GPT-OSS† 24.0 120 / 5.1 73.6 48.0 61.4 61.0
Llama-3.3-Ins 22.6 70 72.1 45.7 61.9 59.9
Qwen3† 20.9 8 70.3 42.9 59.9 57.7
Mistral-Small-3.2 20.7 24 70.6 44.3 62.6 59.2
Qwen3 20.6 32 73.5 43.4 62.8 59.9
RedSage-Ins 19.7 8 67.9 42.4 58.8 56.4
GPT-OSS† 19.6 20 / 3.6 72.3 42.2 57.1 57.2
Gemma-3-it 18.9 27 73.4 40.9 60.2 58.1
Llama-Primus-Ins 18.8 70 72.8 41.9 59.6 58.1
Qwen3 18.3 8 68.9 39.3 58.5 55.5
RedSage-DPO 18.2 8 66.8 40.1 56.7 54.5
Llama-Primus-Merged 16.3 8 66.2 36.7 56.6 53.2
Llama-Primus-Base 15.8 8 67.3 36.6 56.3 53.4
Foundation-Sec-Ins 14.7 8 66.5 33.5 53.0 51.0
Foundation-Sec-Rsn† 13.3 8 67.0 34.8 53.0 51.6
Llama3.1-Ins 12.5 8 61.7 32.4 53.5 49.2

Source: Paper, Table 1. Reported scores are rounded to one decimal place; mode averages include all 24 configurations. This is the paper’s evaluation snapshot, not a live leaderboard.

Our models are named RedSage-K here. They correspond to the paper’s Kali-SFT+GRPO, Kali-SFT, and Kali-GRPO variants, respectively.

How are the scores calculated?
Tool accuracy
Whether the model selects the correct tool.
Optional-argument F1
Alias-aware matching of optional flag names.
Positional-argument F1
Multiset matching of positional argument values.
Total score
The mean of tool accuracy, optional-argument F1, and positional-argument F1.
Exact correct
The paper’s end-to-end metric for canonicalized command correctness. It permits documented aliases and reordered optional flags, while respecting positional argument order.

All scores are percentages; higher is better. Full metric definitions: paper, Appendix L.

Beyond open weightsProprietary models & scaffolded inference

Table 2 evaluates three additional systems in the unrestricted setting. Full-set Exact Correct counts blocked or missing responses as incorrect, making response coverage part of the comparison.

Additional unrestricted results, paper Table 2
System Answered / 5,000 EC, answered EC, all 5,000
GPT-5.6-Sol 4,988 61.83% 61.68%
Codex CLI (GPT-5.5, xhigh) 4,633 55.77% 51.68%
Claude Opus 5 3,675 59.89% 44.02%

Source: Paper, Table 2. Codex is evaluated as a scaffolded CLI system; only its final submitted command is scored.

Dataset construction

Documentation-grounded generation, multiple rounds of verification, and semantic deduplication produce a compact dataset with broad tool coverage.

KaliBench construction pipeline. Official Kali documentation grounds generation; LLM verification, sandbox execution, and human review validate the examples; semantic deduplication produces 8,504 pairs for fine-grained evaluation.
Figure 2. The end-to-end pipeline from the current paper. Open full-size figure.
  1. Generate from documentation

    Tool names, documented flags, values, and aliases ground 27.7K candidate query–command pairs.

  2. Verify and refine

    LLM checks, sandboxed Kali execution, and human review remove unsupported syntax and labeling errors.

  3. Deduplicate and split

    8,504 verified pairs remain: 3,504 for training and 5,000 for evaluation, with all 1,642 tools represented in the test set.

Coverage across 23 dimensions and 5 security phases
  1. 01

    Reconnaissance & Initial Access

    Information gathering · 802.11 · RFID · Bluetooth · SDR · Social engineering

    6
  2. 02

    Vulnerability Analysis

    Vulnerability · Fuzzing · Web · Database · VoIP

    5
  3. 03

    Exploitation & Payload Delivery

    Exploitation · Sniffing & spoofing · Hardware · Wireless

    4
  4. 04

    Post-Exploitation & Lateral Movement

    Post-exploitation · Passwords · Windows resources · Cryptography & steganography

    4
  5. 05

    Defensive Analysis & Reporting

    Reverse engineering · Forensics · Reporting · GPU

    4

Taxonomy: paper, Section 3.

Training with verifiable rewards

Ground-truth commands are verified during dataset construction. During training, deterministic scoring provides rewards for tool selection, arguments, exact command correctness, and output format—without executing model predictions.

Starting from RedSage-Ins, SFT followed by GRPO improves the 8B model’s average total score from 71.7 to 79.2.

Training guide
Post-training comparisonAverage total score (%)
RedSage-Ins 8B · baseline71.7
RedSage-K (SFT) 8B77.4
RedSage-K (SFT+GRPO) 8B · ours79.2
DeepSeek-V3.2 685B / 37B active80.2

Average across all three modes · Paper, Table 1. SFT+GRPO is 1.0 point behind DeepSeek-V3.2. Gains are strongest without documentation; GRPO can reduce hinted-mode performance.

Paper, data, and code

Everything needed to inspect the benchmark, train a model, or reproduce the evaluation.

Run your own evaluation

Clone the repository, set up the environment, and choose a local model or Hugging Face model ID.

Evaluation guide
Get the repository
git clone https://github.com/RISys-Lab/KaliBench.git
cd KaliBench

Environment setup and GPU requirements are included in the guide.

Citation

If you use KaliBench in your research, please cite our paper.

Download BibTeX
BIBTEX
@inproceedings{li2026kalibench,
  title     = {{KaliBench}: A Fine-Grained Benchmark for
               Cybersecurity Tool Use on {Kali Linux}
               with Runtime-Free Verifiable Rewards},
  author    = {Pengfei Li and Naufal Suryanto and Sicheng Zhang and Muzammal Naseer},
  booktitle = {The Fortieth Annual Conference on Neural Information Processing Systems 
               Evaluations and Datasets Track},
  year      = {2026},
  note      = {NeurIPS 2026 Evaluations and Datasets Track},
  url       = {https://github.com/RISys-Lab/KaliBench}
}