From the 1 of 9 linked papers with an AI index.
1 paper · 1 filter
Yanyu Ren, Xizheng Wang, Xiao Liu +8
Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with gro…