1 paper · 1 filter
Jiayu Wang, Yifei Ming, Zixuan Ke +4
Reinforcement learning (RL) has become the dominant paradigm for improving the performance of language models on complex reasoning tasks. Despite the substantial empirical gains de…