1 paper
Yimeng Ye, Shuang Chen, Wenxuan Huang +8
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. Th…