187 citations · 187 across the 4 of their papers we have counts for
1 paper · 1 filter
Wei Gao, Yuheng Zhao, Dakai An +11
Reinforcement Learning (RL) is a pivotal post-training technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, synchronous RL post-training oft…