5 citations · 6 across the 6 of their papers we have counts for
1 paper · 1 filter
Pinzheng Wang, Shuli Xu, Juntao Li +4
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning performance of large language models (LLMs) by increasing test-time compute. Howe…