1 paper · 1 filter
Yixing Chen, Yiding Wang, Siqi Zhu +5
Reinforcement Learning (RL) has demonstrated significant potential in enhancing the reasoning capabilities of large language models (LLMs). However, the success of RL for LLMs heav…