1 paper · 1 filter
Xiao Hu, Xingyu Lu, Liyuan Mao +6
Reinforcement learning (RL) has played an important role in improving the reasoning ability of large language models (LLMs). Some studies apply RL directly to \textit{smaller} base…