1 paper · 1 filter
Haozhe Wang, Qixin Xu, Che Liu +3
Reinforcement Learning (RL) has proven highly effective at enhancing the complex reasoning abilities of Large Language Models (LLMs), yet underlying mechanisms driving this success…