1 paper · 1 filter
Hongzhan Chen, Tao Yang, Yuhua Zhu +3
While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to super…