1 paper
Yuan Li, Bo Wang, Yufei Gao +4
Proximal constraints are fundamental to the stability of the Large Language Model reinforcement learning. While the canonical clipping mechanism in PPO serves as an efficient surro…