1 paper · 1 filter
Zhipeng Chen, Xiaobo Qin, Wayne Xin Zhao +2
Reinforcement learning with verifiable rewards (RLVR) has shown great potential to enhance the reasoning ability of large language models (LLMs). However, due to the limited amount…