1 paper · 1 filter
Zhihe Yang, Xufang Luo, Zilong Wang +4
Reinforcement learning (RL) has become a cornerstone for enhancing the reasoning capabilities of large language models (LLMs), with recent innovations such as Group Relative Policy…