1 citations · 2 across the 5 of their papers we have counts for
1 paper · 1 filter
Tianhao Hu, Xiangcheng Liu, Yuchun Miao +16
Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed ge…