This repository contains the public implementation of Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards (CBS), which introduces a neural scheduler that adaptively selects high-value rollouts during RLVR training, simultaneously addressing noisy intra-group rollout selection and inefficient reuse of historical rollouts. The code is built on top of verl and provides the implementation of the CBS and CBS_star rollout schedulers, together with ready-to-run example recipes.
CBS is based on verl 0.5.x. Create a Python 3.10.12 environment first:
conda create -n cbs python=3.10.12 -y
conda activate cbsThen follow the official verl v0.5.x installation guide to install the base training packages. Please make sure the environment matches the core dependency versions used by our experiments:
Python 3.10.12
CUDA 12.4
PyTorch 2.6.0
Ray 2.43.0
Transformers 4.51.3
vLLM 0.8.5
xFormers 0.0.29.post2
Additionally install the math-verification packages required by the reward functions:
pip install math_verify
pip install pylatexencThe repository already includes the processed parquet files used by the example recipes, so the scripts can be run directly without rebuilding the datasets from raw sources. The default training files are:
data/dapo/train.parquet
data/math/train.parquet
The default validation files used during training are:
data/aime25/test_16.parquet
data/aime24/test_16.parquet
data/amc23/test_16.parquet
data/math500/test.parquet
data/minerva/test.parquet
data/olympiad/test.parquet
We also provide the larger evaluation splits used by the evaluation scripts:
data/aime25/test_32.parquet
data/aime24/test_32.parquet
data/amc23/test_32.parquet
data/math500/test_4.parquet
data/minerva/test_4.parquet
data/olympiad/test_4.parquet
If your data or model checkpoints are stored elsewhere, edit train_files, test_files, and model_path at the top of the corresponding script in recipes/.
The example scripts provide ready-to-run commands for the baseline, CBS, and CBS_star on Qwen3-1.7B-Base, Qwen3-4B-Base, and Qwen3-8B-Base. Each script is configured for a single 8-GPU node by default and writes logs, rollouts, validation outputs, and checkpoints under outputs/, tensorboard_log/, and checkpoints/.
## Qwen3-1.7B-Base
bash recipes/Qwen3_1.7B_Base_GRPO.sh
bash recipes/Qwen3_1.7B_Base_GRPO_CBS.sh
bash recipes/Qwen3_1.7B_Base_GRPO_CBS_star.sh
## Qwen3-4B-Base
bash recipes/Qwen3_4B_Base_GRPO.sh
bash recipes/Qwen3_4B_Base_GRPO_CBS.sh
bash recipes/Qwen3_4B_Base_GRPO_CBS_star.sh
## Qwen3-8B-Base
bash recipes/Qwen3_8B_Base_GRPO.sh
bash recipes/Qwen3_8B_Base_GRPO_CBS.sh
bash recipes/Qwen3_8B_Base_GRPO_CBS_star.shIf you find this repository useful, please cite our paper:
@inproceedings{lu2026contextual,
title = {Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards},
author = {Lu, Xiaodong and Wang, Xiaohan and Chai, Jiajun and Yin, Guojun and Lin, Wei and Chen, Zhijun and Luo, Yu and Zhuang, Fuzhen and Ban, Yikun and Wang, Deqing},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
year = {2026}
}