106 citations · 128 across the 8 of their papers we have counts for
1 paper · 2 filters
Jie Xiao, Changyuan Fan, Qingnan Ren +6
Modern RL-based post-training for large language models (LLMs) co-locate trajectory sampling and policy optimisation on the same GPU cluster, forcing the system to switch between i…