1 paper · 1 filter
Naoki Shitanda, Motoki Omura, Tatsuya Harada +1
Scaling reinforcement learning to tens of thousands of parallel environments requires overcoming the limited exploration capacity of a single policy. Ensemble-based policy gradient…