89 citations · 243 across the 36 of their papers we have counts for
1 paper · 2 filters
Hoang Phan, Xianjun Yang, Yuanshun Yao +6
Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for c…