1 paper
Hanqing Zhu, Wenyan Cong, Zhizhou Sha +8
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback i…