2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Shihui Yang, Chengfeng Dou, Peidong Guo +4
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning capabilities of large language models. However, existing appr…