Evaluation and scoring

Environment setup

Run these commands from a clone of the repository. Python 3.10 is recommended.

git clone https://github.com/RISys-Lab/KaliBench.git
cd KaliBench
conda create -n kalibench python=3.10
conda activate kalibench
pip install -r requirements.txt

Evaluation uses vLLM. PyTorch, CUDA, vLLM, and FlashInfer/Triton builds must be compatible with your GPU and driver. Model downloads require network access unless already cached.

Complete the environment setup and run all commands from the repository root. The examples use the merged model from the training guide; replace --model with another local model path or Hugging Face model ID as needed.

src/evaluate.py builds prompts and labels, runs vLLM inference, checkpoints predictions, and calculates tool, optional F1, positional F1, total, exact-match, and format-error metrics. Scoring across the 23 Kali tool dimensions is optional.

Use a different output directory for every model/mode pair. This prevents a resumed run from mixing predictions generated with different prompts.

Evaluate one mode

python src/evaluate.py \
  --mode hinted \
  --input "$PWD/KaliBench_data/kalibench_verified_test_5000.jsonl" \
  --subtools "$PWD/KaliBench_data/Kali_Tool_Subtools_UsageCode.jsonl" \
  --model "$PWD/outputs/models/kalibench_grpo/merged" \
  --output-dir "$PWD/outputs/evaluate/kalibench_grpo/hinted" \
  --candidate-seed 42 \
  --seed 3407 \
  --temperature 0.2 \
  --batch-size 128 \
  --tensor-parallel-size 1

In hinted mode, evaluation intentionally forces one candidate tool and includes its usage information. In restricted mode, --candidate-tools 20 is the default. Unrestricted mode does not use a candidate-tool list.

Evaluate all three modes

for MODE in unrestricted restricted hinted; do
  python src/evaluate.py \
    --mode "$MODE" \
    --input "$PWD/KaliBench_data/kalibench_verified_test_5000.jsonl" \
    --subtools "$PWD/KaliBench_data/Kali_Tool_Subtools_UsageCode.jsonl" \
    --model "$PWD/outputs/models/kalibench_grpo/merged" \
    --output-dir "$PWD/outputs/evaluate/kalibench_grpo/$MODE" \
    --candidate-tools 20 \
    --candidate-seed 42 \
    --seed 3407 \
    --temperature 0.2 \
    --batch-size 128 \
    --tensor-parallel-size 1
done

Resume and runtime options

The default --resume true behavior reads existing raw_predictions.jsonl and vllm_compat_predictions.jsonl files, validates their IDs, and evaluates only missing examples. Use --resume false to start a clean run in an existing output directory.

Useful hardware/runtime flags include:

--tensor-parallel-size
--gpu-memory-utilization
--max-model-len
--max-num-batched-tokens
--dtype
--enforce-eager
--language-model-only
--gdn-prefill-backend
--inference-api

Run python src/evaluate.py --help for their complete descriptions.

Evaluation outputs

Each evaluation directory contains:

run_config.json
labels.jsonl
raw_predictions.jsonl
vllm_compat_predictions.jsonl
<model>.<mode>.summary.json
model_scores/<model>.<mode>.scores.jsonl

--output-raw additionally stores the actual prompt messages and input/output token counts. --wandb enables optional run configuration and summary logging.

Scoring and result generation

Re-score saved predictions

Scoring runs automatically unless --skip-scoring is passed. To re-score an existing evaluation without repeating inference:

python src/scoring_pipeline.py \
  --folder "$PWD/outputs/evaluate/kalibench_grpo/hinted"

For explicit prediction and label files:

python src/model_score.py \
  --vllm_file "$PWD/outputs/evaluate/kalibench_grpo/hinted/vllm_compat_predictions.jsonl" \
  --label_file "$PWD/outputs/evaluate/kalibench_grpo/hinted/labels.jsonl" \
  --out "$PWD/outputs/evaluate/kalibench_grpo/hinted/model_scores/rescored.jsonl"

Dimension-level scores

Dimension-level scoring uses the tool-to-title mapping in the released usage file and the anonymous Hugging Face metadata for the 23 metapackage labels. This lookup requires network access unless the metadata is already cached:

python Scoring/dim_score.py \
  --models-folder "$PWD/outputs/evaluate/kalibench_grpo/hinted/model_scores" \
  --output-folder "$PWD/outputs/evaluate/kalibench_grpo/hinted/dim_scores" \
  --subtools "$PWD/KaliBench_data/Kali_Tool_Subtools_UsageCode.jsonl" \
  --use-hf \
  --hf-dataset "anonymous62567/kali-tools"

The same step can be appended to end-to-end evaluation with:

--run-dim-score --dim-use-hf --dim-hf-dataset anonymous62567/kali-tools

Paper-style result tables

src/format_table_results.py converts an aggregate run CSV into an organized CSV and a LaTeX table grouped by model size. Its input expects one row per model/mode configuration with config_* and summary_* columns.

python src/format_table_results.py aggregate_results.csv \
  --output-csv organized_scores.csv \
  --output-tex organized_scores.tex \
  --size-grouping coarse

Use --override-sizes MODEL:7B for model names that do not contain a recognizable parameter count.

Adapted from the repository’s docs/evaluation.md. Commands are shown for reference; this page does not run them.