1 paper
Tianhao Hu, Xiangcheng Liu, Youshao Xiao +21
Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed ge…