1 paper
Weisong Zhao, Tong Wang, Zichang Tan +11
Group-based reinforcement learning has evolved from the arithmetic mean of GRPO to the geometric mean of GMPO. While GMPO improves stability by constraining a conservative objectiv…