1 paper
Lucas Hu, Ranchi Zhao, Isaac Zhu +4
In large-scale reinforcement learning (RL) systems with decoupled Trainer-Rollout execution, the Trainer must regularly synchronize policy weights to the Rollout side to limit poli…