1 paper
Youngjun Yu, Sanghwan Jang, Hwanjo Yu
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample acc…