1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Kaiyi Zhang, Ang Lv, Jinpeng Li +4
Reinforcement learning with verifiable rewards (RLVR) is a promising approach for improving the complex reasoning abilities of large language models (LLMs). However, current RLVR m…