A Workflow for Offline Model-Free Robotic Reinforcement Learning
arXiv:2109.10813
Abstract
Offline reinforcement learning (RL) enables learning control policies by utilizing only prior experience, without any online interaction. This can allow robots to acquire generalizable skills from large and diverse datasets, without any costly or unsafe online data collection. Despite recent algorithmic advances in offline RL, applying these methods to real-world problems has proven challenging. Although offline RL methods can learn from prior data, there is no clear and well-understood process for making various design choices, from model architecture to algorithm hyperparameters, without actually evaluating the learned policies online. In this paper, our aim is to develop a practical workflow for using offline RL analogous to the relatively well-understood workflows for supervised learning problems. To this end, we devise a set of metrics and conditions that can be tracked over the course of offline training, and can inform the practitioner about how the algorithm and model architecture should be adjusted to improve final performance. Our workflow is derived from a conceptual understanding of the behavior of conservative offline RL algorithms and cross-validation in supervised learning. We demonstrate the efficacy of this workflow in producing effective policies without any online tuning, both in several simulated robotic learning scenarios and for three tasks on two distinct real robots, focusing on learning manipulation skills with raw image observations with sparse binary rewards. Explanatory video and additional results can be found at sites.google.com/view/offline-rl-workflow
CoRL 2021. Project Website: https://sites.google.com/view/offline-rl-workflow. First two authors contributed equally
References in corpus (25)
- Continuous control with deep reinforcement learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Conservative Q-Learning for Offline Reinforcement Learning
- Deep Variational Information Bottleneck
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Behavior Regularized Offline Reinforcement Learning
- MOPO: Model-based Offline Policy Optimization
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- A Minimalist Approach to Offline Reinforcement Learning
- MOReL : Model-Based Offline Reinforcement Learning
- Collective Robot Reinforcement Learning with Distributed Asynchronous Guided Policy Search
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning
- Learning Dexterous Manipulation Policies from Experience and Imitation
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
- Reinforcement Learning via Fenchel-Rockafellar Duality
- NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning
- Visual Imitation Made Easy
- Interpretable Latent Spaces for Learning from Demonstration
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills
- Improvisation through Physical Understanding: Using Novel Objects as Tools with Visual Foresight
- Statistical Bootstrapping for Uncertainty Estimation in Off-Policy Evaluation
- Continuous Doubly Constrained Batch Reinforcement Learning