1 paper
Xikai Zhang, Yongzhi Li, Likang Xiao +6
Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants…