34 citations · 34 across the 3 of their papers we have counts for
1 paper · 1 filter
Xikai Zhang, Yongzhi Li, Likang Xiao +6
Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants…