1 paper · 1 filter
Ryan Xu, Atlas Zhao, David Bao +1
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectorie…