2 citations · 4 across the 15 of their papers we have counts for
1 paper · 1 filter
Zhihe Yang, Xufang Luo, Zilong Wang +4
Reinforcement learning (RL) has become a cornerstone for enhancing the reasoning capabilities of large language models (LLMs), with recent innovations such as Group Relative Policy…