1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Zhipeng Chen, Xiaobo Qin, Youbin Wu +4
Reinforcement learning with verifiable rewards (RLVR), which typically adopts Pass@1 as the reward, has faced the issues in balancing exploration and exploitation, causing policies…