1 paper
Naoki Shitanda, Motoki Omura, Tatsuya Harada +1
Scaling reinforcement learning to tens of thousands of parallel environments requires overcoming the limited exploration capacity of a single policy. Ensemble-based policy gradient…