Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
arXiv:2109.08128
Abstract
Offline reinforcement learning (RL) algorithms have shown promising results in domains where abundant pre-collected data is available. However, prior methods focus on solving individual problems from scratch with an offline dataset without considering how an offline RL agent can acquire multiple skills. We argue that a natural use case of offline RL is in settings where we can pool large amounts of data collected in various scenarios for solving different tasks, and utilize all of this data to learn behaviors for all the tasks more effectively rather than training each one in isolation. However, sharing data across all tasks in multi-task offline RL performs surprisingly poorly in practice. Thorough empirical analysis, we find that sharing data can actually exacerbate the distributional shift between the learned policy and the dataset, which in turn can lead to divergence of the learned policy and poor performance. To address this challenge, we develop a simple technique for data-sharing in multi-task offline RL that routes data based on the improvement over the task-specific data. We call this approach conservative data sharing (CDS), and it can be applied with multiple single-task offline RL methods. On a range of challenging multi-task locomotion, navigation, and vision-based robotic manipulation problems, CDS achieves the best or comparable performance compared to prior offline multi-task RL methods and previous data sharing approaches.
References in corpus (51)
- Continuous control with deep reinforcement learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Trust Region Policy Optimization
- Addressing Function Approximation Error in Actor-Critic Methods
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Conservative Q-Learning for Offline Reinforcement Learning
- Deep reinforcement learning from human preferences
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- Hindsight Experience Replay
- Visual Reinforcement Learning with Imagined Goals
- Behavior Regularized Offline Reinforcement Learning
- MOPO: Model-based Offline Policy Optimization
- Maximum a Posteriori Policy Optimisation
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- Distral: Robust Multitask Reinforcement Learning
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- MOReL : Model-Based Offline Reinforcement Learning
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Gradient Surgery for Multi-Task Learning
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Safe Policy Improvement with Baseline Bootstrapping
- Sharing Knowledge in Multi-Task Deep Reinforcement Learning
- Multi-Task Reinforcement Learning with Soft Modularization
- Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
- COMBO: Conservative Offline Model-Based Policy Optimization
- Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
- Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning
- Off-Policy Policy Gradient with State Distribution Correction
- Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning
- Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation
- Provably Good Batch Reinforcement Learning Without Great Exploration
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
- Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Offline Reinforcement Learning with Fisher Divergence Critic Regularization
- Representation Matters: Offline Pretraining for Sequential Decision Making
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous Control
- Multi-Task Reinforcement Learning with Context-based Representations
- Model-Based Offline Planning
- Generalized Hindsight for Reinforcement Learning
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills
- Improvisation through Physical Understanding: Using Novel Objects as Tools with Visual Foresight
- PLAS: Latent Action Space for Offline Reinforcement Learning
- An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare
- Multi-task Batch Reinforcement Learning with Metric Learning
- Divide-and-Conquer Reinforcement Learning
- Reinforcement Learning without Ground-Truth State
- Competitive Experience Replay
- Bias-reduced Multi-step Hindsight Experience Replay for Efficient Multi-goal Reinforcement Learning