Observe and Look Further: Achieving Consistent Performance on Atari
arXiv:1805.11593
Abstract
Despite significant advances in the field of deep Reinforcement Learning (RL), today's algorithms still fail to learn human-level policies consistently over a set of diverse tasks such as Atari 2600 games. We identify three key challenges that any algorithm needs to master in order to perform well on all games: processing diverse reward distributions, reasoning over long time horizons, and exploring efficiently. In this paper, we propose an algorithm that addresses each of these challenges and is able to learn human-level policies on nearly all Atari games. A new transformed Bellman operator allows our algorithm to process rewards of varying densities and scales; an auxiliary temporal consistency loss allows us to train stably using a discount factor of (instead of ) extending the effective planning horizon by an order of magnitude; and we ease the exploration problem by using human demonstrations that guide the agent towards rewarding states. When tested on a set of 42 Atari games, our algorithm exceeds the performance of an average human on 40 games using a common set of hyper parameters. Furthermore, it is the first deep RL algorithm to solve the first level of Montezuma's Revenge.
References in corpus (5)
Cited by in corpus (40)
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
- Exploration by Random Network Distillation
- Agent57: Outperforming the Atari Human Benchmark
- Go-Explore: a New Approach for Hard-Exploration Problems
- Ablation Studies in Artificial Neural Networks
- Sample Efficient Adaptive Text-to-Speech
- Online Scheduling of a Residential Microgrid via Monte-Carlo Tree Search and a Learned Model
- Learning Montezuma's Revenge from a Single Demonstration
- Never Give Up: Learning Directed Exploration Strategies
- Towards Characterizing Divergence in Deep Q-Learning
- Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
- Scaling data-driven robotics with reward sketching and batch reinforcement learning
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- Is Deep Reinforcement Learning Really Superhuman on Atari? Leveling the playing field
- State Alignment-based Imitation Learning
- Deep Q-learning: a robust control approach
- Offline Learning from Demonstrations and Unlabeled Experience
- Reinforcement Learning with A* and a Deep Heuristic
- Muesli: Combining Improvements in Policy Optimization
- World Discovery Models
- Using a Logarithmic Mapping to Enable Lower Discount Factors in Reinforcement Learning
- General non-linear Bellman equations
- Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments
- ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations
- Return-based Scaling: Yet Another Normalisation Trick for Deep RL
- I'm sorry Dave, I'm afraid I can't do that, Deep Q-learning from forbidden action
- Expert-augmented actor-critic for ViZDoom and Montezumas Revenge
- Convergent and Efficient Deep Q Network Algorithm
- Learning and Exploiting Multiple Subgoals for Fast Exploration in Hierarchical Reinforcement Learning
- Regularized Behavior Value Estimation
- Taylor Expansion Policy Optimization
- Learning Abstract Models for Strategic Exploration and Fast Reward Transfer
- Continuous Control for Searching and Planning with a Learned Model
- From semantics to execution: Integrating action planning with reinforcement learning for robotic causal problem-solving
- ConQUR: Mitigating Delusional Bias in Deep Q-learning
- Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation
- Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy
- DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning
- Evolving Indoor Navigational Strategies Using Gated Recurrent Units In NEAT
- Self-Consistent Models and Values