Rainbow: Combining Improvements in Deep Reinforcement Learning
arXiv:1710.02298
Abstract
The deep reinforcement learning community has made several independent improvements to the DQN algorithm. However, it is unclear which of these extensions are complementary and can be fruitfully combined. This paper examines six extensions to the DQN algorithm and empirically studies their combination. Our experiments show that the combination provides state-of-the-art performance on the Atari 2600 benchmark, both in terms of data efficiency and final performance. We also provide results from a detailed ablation study that shows the contribution of each component to overall performance.
Under review as a conference paper at AAAI 2018
Cited by in corpus (16)
- Hyperbolic Discounting and Learning over Multiple Horizons
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
- Hyp-RL : Hyperparameter Optimization by Reinforcement Learning
- A Comparative Analysis of Expected and Distributional Reinforcement Learning
- Experience Replay Optimization
- Modern Deep Reinforcement Learning Algorithms
- Learning to Navigate in Indoor Environments: from Memorizing to Reasoning
- World Discovery Models
- General non-linear Bellman equations
- DOM-Q-NET: Grounded RL on Structured Language
- MULEX: Disentangling Exploitation from Exploration in Deep RL
- AgentGraph: Towards Universal Dialogue Management with Structured Deep Reinforcement Learning
- Finding Needles in a Moving Haystack: Prioritizing Alerts with Adversarial Reinforcement Learning
- Learning Gaussian Policies from Corrective Human Feedback
- Growing Action Spaces