1 paper
Ryan Xu, Atlas Zhao, David Bao +1
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectorie…