A Survey on Offline Reinforcement Learning: Taxonomy, Review, and Open Problems
arXiv:2203.01387 · doi:10.1109/TNNLS.2023.3250269
Abstract
With the widespread adoption of deep learning, reinforcement learning (RL) has experienced a dramatic increase in popularity, scaling to previously intractable problems, such as playing complex games from pixel observations, sustaining conversations with humans, and controlling robotic agents. However, there is still a wide range of domains inaccessible to RL due to the high cost and danger of interacting with the environment. Offline RL is a paradigm that learns exclusively from static datasets of previously collected interactions, making it feasible to extract policies from large and diverse training datasets. Effective offline RL algorithms have a much wider range of applications than online RL, being particularly appealing for real-world applications, such as education, healthcare, and robotics. In this work, we contribute with a unifying taxonomy to classify offline RL methods. Furthermore, we provide a comprehensive review of the latest algorithmic breakthroughs in the field using a unified notation as well as a review of existing benchmarks' properties and shortcomings. Additionally, we provide a figure that summarizes the performance of each method and class of methods on different dataset properties, equipping researchers with the tools to decide which type of algorithm is best suited for the problem at hand and identify which classes of algorithms look the most promising. Finally, we provide our perspective on open problems and propose future research directions for this rapidly growing field.
21 pages; Final version accepted to IEEE Transactions on Neural Networks and Learning Systems
References in corpus (13)
- A Brief Survey of Deep Reinforcement Learning
- DeepMind Control Suite
- Behavior Regularized Offline Reinforcement Learning
- dm_control: Software and Tasks for Continuous Control
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Offline Reinforcement Learning with Implicit Q-Learning
- AlphaStar: An Evolutionary Computation Perspective
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning
- Hyperparameter Selection for Offline Reinforcement Learning
- Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning
- Reinforcement Learning for Education: Opportunities and Challenges
Cited by in corpus (14)
- Autonomous Navigation for Robot-assisted Intraluminal and Endovascular Procedures: A Systematic Review
- A Review of Safe Reinforcement Learning Methods for Modern Power Systems
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- A Safe Deep Reinforcement Learning Approach for Energy Efficient Federated Learning in Wireless Communication Networks
- ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning
- AD4RL: Autonomous Driving Benchmarks for Offline Reinforcement Learning with Value-based Dataset
- ORL-AUDITOR: Dataset Auditing in Offline Deep Reinforcement Learning
- Framework for Learning and Control in the Classical and Quantum Domains
- Unsupervised Meta-Testing with Conditional Neural Processes for Hybrid Meta-Reinforcement Learning
- A Model-Based Approach for Improving Reinforcement Learning Efficiency Leveraging Expert Observations
- A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning
- Learning from Imperfect Demonstrations with Self-Supervision for Robotic Manipulation
- Meta-Offline and Distributional Multi-Agent RL for Risk-Aware Decision-Making
- Enhancing Reinforcement Learning Through Guided Search