Large Batch Experience Replay
arXiv:2110.01528
Abstract
Several algorithms have been proposed to sample non-uniformly the replay buffer of deep Reinforcement Learning (RL) agents to speed-up learning, but very few theoretical foundations of these sampling schemes have been provided. Among others, Prioritized Experience Replay appears as a hyperparameter sensitive heuristic, even though it can provide good performance. In this work, we cast the replay buffer sampling problem as an importance sampling one for estimating the gradient. This allows deriving the theoretically optimal sampling distribution, yielding the best theoretical convergence speed. Elaborating on the knowledge of the ideal sampling scheme, we exhibit new theoretical foundations of Prioritized Experience Replay. The optimal sampling distribution being intractable, we make several approximations providing good results in practice and introduce, among others, LaBER (Large Batch Experience Replay), an easy-to-code and efficient method for sampling the replay buffer. LaBER, which can be combined with Deep Q-Networks, distributional RL agents or actor-critic methods, yields improved performance over a diverse range of Atari games and PyBullet environments, compared to the base agent it is implemented on and to other prioritization schemes.
24 pages, 12 figures, ICML 2022 - long presentation
References in corpus (11)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Addressing Function Approximation Error in Actor-Critic Methods
- Distributed Prioritized Experience Replay
- Implicit Quantile Networks for Distributional Reinforcement Learning
- A Deeper Look at Experience Replay
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Not All Samples Are Created Equal: Deep Learning with Importance Sampling
- MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- Revisiting Rainbow: Promoting more Insightful and Inclusive Deep Reinforcement Learning Research
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience Replay