1 paper
Zhaopeng Feng, Shaosheng Cao, Jiahan Ren +7
Large-scale reinforcement learning (RL) methods have proven highly effective in enhancing the reasoning abilities of large language models (LLMs), particularly for tasks with verif…