This repository contains the official implementation of the paper Reinforced Efficient Reasoning via Semantically Diverse Exploration. This project is built on top of VERL, a framework for reinforcement learning on large language models.
python examples/data_preprocess/preprocess_all_math.pyTo run an example training script, execute:
./example.shSee requirements.txt for Python dependencies.
Citation
If you find our work useful, please cite our paper:
@article{zhao2026reer-sde,
title={Reinforced Efficient Reasoning via Semantically Diverse Exploration},
author={Ziqi Zhao and Zhaochun Ren and Jiahong Zou and Liu Yang and Zhiwei Xu and Xuri Ge and Zhumin Chen and Xinyu Ma and Daiting Shi and Shuaiqiang Wang and Dawei Yin and Xin Xin},
journal={arXiv preprint arXiv:2601.05053},
year={2026},
url={https://arxiv.org/abs/2601.05053}
}