1 paper · 1 filter
Jianren Wang, Yifan Su, Abhinav Gupta +1
On-policy reinforcement learning (RL) algorithms are widely used for their strong asymptotic performance and training stability, but they struggle to scale with larger batch sizes,…