1 citations · 2 across the 8 of their papers we have counts for
1 paper · 1 filter
Huimin Xu, Shuai Zhao, Xiaobao Wu +1
Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning ability of large language models. However, widely used RLVR algor…