An Optimistic Perspective on Offline Reinforcement Learning
arXiv:1907.04543
Abstract
Off-policy reinforcement learning (RL) using a fixed offline dataset of logged interactions is an important consideration in real world applications. This paper studies offline RL using the DQN replay dataset comprising the entire replay experience of a DQN agent on 60 Atari 2600 games. We demonstrate that recent off-policy deep RL algorithms, even when trained solely on this fixed dataset, outperform the fully trained DQN agent. To enhance generalization in the offline setting, we present Random Ensemble Mixture (REM), a robust Q-learning algorithm that enforces optimal Bellman consistency on random convex combinations of multiple Q-value estimates. Offline REM trained on the DQN replay dataset surpasses strong RL baselines. Ablation studies highlight the role of offline dataset size and diversity as well as the algorithm choice in our positive results. Overall, the results here present an optimistic view that robust RL algorithms trained on sufficiently large and diverse offline datasets can lead to high quality policies. The DQN replay dataset can serve as an offline RL benchmark and is open-sourced.
ICML 2020. An earlier version was titled "Striving for Simplicity in Off-Policy Deep Reinforcement Learning". Project Website: https://offline-rl.github.io
References in corpus (10)
- Model-Based Reinforcement Learning for Atari
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Challenges of Real-World Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Learning from Logged Implicit Exploration Data
- When to use parametric models in reinforcement learning?
- Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning
Cited by in corpus (14)
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Conservative Q-Learning for Offline Reinforcement Learning
- Reinforcement Learning for Selective Key Applications in Power Systems: Recent Advances and Future Challenges
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive Recommendation
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning
- Unifying Cardiovascular Modelling with Deep Reinforcement Learning for Uncertainty Aware Control of Sepsis Treatment
- Value Penalized Q-Learning for Recommender Systems
- Why People Skip Music? On Predicting Music Skips using Deep Reinforcement Learning
- ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning
- Regularized Behavior Value Estimation
- OER: Offline Experience Replay for Continual Offline Reinforcement Learning
- Accelerating Offline Reinforcement Learning Application in Real-Time Bidding and Recommendation: Potential Use of Simulation
- Few-Shot Image-to-Semantics Translation for Policy Transfer in Reinforcement Learning
- The Least Restriction for Offline Reinforcement Learning