Training
Environment setup
Run these commands from a clone of the repository. Python 3.10 is recommended.
git clone https://github.com/RISys-Lab/KaliBench.git
cd KaliBench
conda create -n kalibench python=3.10
conda activate kalibench
pip install -r requirements.txt
Evaluation uses vLLM. PyTorch, CUDA, vLLM, and FlashInfer/Triton builds must be compatible with your GPU and driver. Model downloads require network access unless already cached.
Complete the environment setup, then install the training dependencies:
pip install unsloth trl wandb
Unsloth, TRL, PyTorch, and CUDA must be compatible with your GPU and
driver. The examples use --report-to none to disable
Weights & Biases logging. Run all commands from the repository
root.
Set the base model to a local path or a Hugging Face model ID:
export BASE_MODEL="/path/to/base-model"
The main paper-reproduction path is:
constructed and verified training split → KaliBench SFT → GRPO/RLVR → merged model → evaluation
1. Supervised fine-tuning
The KaliBench SFT implementation creates examples in all three modes, applies output-only supervision, trains LoRA adapters with Unsloth/TRL, and saves both the adapter and a merged 16-bit model.
python src/train/sft_kalibench.py \
--model-name "$BASE_MODEL" \
--dataset-path "$PWD/KaliBench_data/kalibench_verified_train_3504.jsonl" \
--subtools-path "$PWD/KaliBench_data/Kali_Tool_Subtools_UsageCode.jsonl" \
--mode hinted restricted unrestricted \
--candidate-tools 20 \
--candidate-seed 42 \
--seed 3407 \
--output-dir "$PWD/outputs/models/kalibench_sft" \
--adapter-output-dir "$PWD/outputs/adapters/kalibench_sft" \
--report-to none
Outputs:
outputs/adapters/kalibench_sft/ # LoRA adapter and tokenizer
outputs/models/kalibench_sft/merged/ # merged 16-bit model
The default optimizer, LoRA rank, sequence length, batch size,
accumulation, epoch count, learning rate, warmup, and checkpoint
cadence are exposed as CLI flags. Run
python src/train/sft_kalibench.py --help for the complete
configuration.
2. GRPO/RLVR
For the SFT+GRPO experiment, initialize GRPO from the merged SFT model. The training objective combines output-format rewards with the same tool, optional-argument, positional-argument, and exact-match signals used during evaluation.
python src/train/grpo_kalibench.py \
--model-name "$PWD/outputs/models/kalibench_sft/merged" \
--dataset-path "$PWD/KaliBench_data/kalibench_verified_train_3504.jsonl" \
--subtools-path "$PWD/KaliBench_data/Kali_Tool_Subtools_UsageCode.jsonl" \
--mode hinted:1.0 restricted:1.0 unrestricted:1.0 \
--candidate-tools 20 \
--candidate-seed 42 \
--seed 3407 \
--output-dir "$PWD/outputs/models/kalibench_grpo" \
--adapter-output-dir "$PWD/outputs/adapters/kalibench_grpo" \
--report-to none
Each --mode value can include a sampling fraction in
[0,1], such as restricted:0.5. The default
reward weights are:
| Reward component | Weight |
|---|---|
| exact output format | 2.0 |
| approximate output format | 1.0 |
| tool score | 1.0 |
| optional-argument F1 | 1.5 |
| positional-argument F1 | 1.5 |
| exact match | 2.0 |
Before launching a long run, the prepared prompts and token-length statistics can be inspected without training:
python src/train/grpo_kalibench.py \
--model-name "$PWD/outputs/models/kalibench_sft/merged" \
--dataset-path "$PWD/KaliBench_data/kalibench_verified_train_3504.jsonl" \
--subtools-path "$PWD/KaliBench_data/Kali_Tool_Subtools_UsageCode.jsonl" \
--mode hinted:1.0 restricted:1.0 unrestricted:1.0 \
--debug-dataset \
--report-to none
3. Optional training utilities
src/train/sft_general.py trains on a generic chat JSONL
file in which every row contains a messages list:
{"messages":[{"role":"system","content":"..."},{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
python src/train/sft_general.py \
--model-name "$BASE_MODEL" \
--dataset-path "/absolute/path/to/messages.jsonl" \
--output-dir "$PWD/outputs/models/general_sft" \
--adapter-output-dir "$PWD/outputs/adapters/general_sft" \
--report-to none
The SFT and GRPO scripts already save a merged model. To merge or export a separately saved adapter:
python src/train/merge_lora.py \
--base-model "$BASE_MODEL" \
--adapter "$PWD/outputs/adapters/kalibench_sft" \
--output-dir "$PWD/outputs/models/kalibench_sft_merged"
Continue with evaluation and scoring after training.
Adapted from the repository’s docs/training.md. Commands
are shown for reference; this page does not run them.