1 paper
Yihua Zhu, Qianying Liu, Fei Cheng +4
Reinforcement learning with verifiable rewards (RLVR) has become central to post-training reasoning models, yet a key limitation of existing studies is their narrow view of the rea…