28 citations · 30 across the 11 of their papers we have counts for
1 paper · 1 filter
Zhiyuan Hu, Yucheng Wang, Yufei He +7
Reinforcement learning (RL) has become a central paradigm for post-training large language models (LLMs), particularly for complex reasoning tasks, yet it often suffers from explor…