1 citations · 2 across the 10 of their papers we have counts for
1 paper · 1 filter
Yuxian Jiang, Yafu Li, Guanxu Chen +3
Reinforcement learning with verifiable rewards (RLVR) has shown great promise in enhancing the reasoning abilities of large reasoning models (LRMs). However, it suffers from a crit…