1 paper · 1 filter
Chuzhan Hao, Wenfeng Feng, Guochao Jiang +3
Reinforcement learning (RL) has become an effective approach for advancing the reasoning capabilities of large language models (LLMs) through the strategic integration of external…