44 citations · 114 across the 22 of their papers we have counts for
1 paper · 1 filter
Hoang Phan, Xianjun Yang, Yuanshun Yao +6
Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for c…