1 paper
Hangfan Zhang, Siyuan Xu, Zhimeng Guo +8
Reinforcement learning (RL) has demonstrated potential in enhancing the reasoning capabilities of large language models (LLMs), but such training typically demands substantial effo…