From the 1 of 14 linked papers with an AI index.
1 paper · 1 filter
Sining Zhoubian, Dan Zhang, Jie Tang
With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method GRPO faces failure due to insignificant reward variance, while verif…