Distral: Robust Multitask Reinforcement Learning
arXiv:1707.04175
Abstract
Most deep reinforcement learning algorithms are data inefficient in complex and rich environments, limiting their applicability to many scenarios. One direction for improving data efficiency is multitask learning with shared neural network parameters, where efficiency may be improved through transfer across related tasks. In practice, however, this is not usually observed, because gradients from different tasks can interfere negatively, making learning unstable and sometimes even less data efficient. Another issue is the different reward schemes between tasks, which can easily lead to one task dominating the learning of a shared model. We propose a new approach for joint training of multiple tasks, which we refer to as Distral (Distill & transfer learning). Instead of sharing parameters between the different workers, we propose to share a "distilled" policy that captures common behaviour across tasks. Each worker is trained to solve its own task while constrained to stay close to the shared policy, while the shared policy is trained by distillation to be the centroid of all task policies. Both aspects of the learning process are derived by optimizing a joint objective function. We show that our approach supports efficient transfer on complex 3D environments, outperforming several related methods. Moreover, the proposed learning process is more robust and more stable---attributes that are critical in deep reinforcement learning.
References in corpus (5)
Cited by in corpus (44)
- A Brief Survey of Deep Reinforcement Learning
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- Deep Reinforcement Learning: An Overview
- Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents
- Multi-Agent Reinforcement Learning via Double Averaging Primal-Dual Optimization
- InfoBot: Transfer and Exploration via the Information Bottleneck
- Robust Reinforcement Learning for Continuous Control with Model Misspecification
- CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
- Kickstarting Deep Reinforcement Learning
- A Neural Entity Coreference Resolution Review
- Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement
- Invariant Causal Prediction for Block MDPs
- Diff-DAC: Distributed Actor-Critic for Average Multitask Deep Reinforcement Learning
- AutoLoss: Learning Discrete Schedules for Alternate Optimization
- Mix&Match - Agent Curricula for Reinforcement Learning
- Pseudo-task Augmentation: From Deep Multitask Learning to Intratask Sharing---and Back
- Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial Networks
- Multi-task Deep Reinforcement Learning with PopArt
- Universal Successor Features Approximators
- Progressive Reinforcement Learning with Distillation for Multi-Skilled Motion Control
- Reinforcement Learning with Perturbed Rewards
- Not to Cry Wolf: Distantly Supervised Multitask Learning in Critical Care
- Self-Attentional Credit Assignment for Transfer in Reinforcement Learning
- Multi-task Learning for Continuous Control
- Learning to Synthesize Programs as Interpretable and Generalizable Policies
- Flatland: a Lightweight First-Person 2-D Environment for Reinforcement Learning
- Fast Adaptation via Policy-Dynamics Value Functions
- Learning Adaptive Exploration Strategies in Dynamic Environments Through Informed Policy Regularization
- Evolutionary Architecture Search For Deep Multitask Networks
- Data-efficient Hindsight Off-policy Option Learning
- Reinforcement Learning for Adaptive Mesh Refinement
- Learning Discrete State Abstractions With Deep Variational Inference
- Meta Automatic Curriculum Learning
- Transfer Learning to Learn with Multitask Neural Model Search
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
- Discovering Fatigued Movements for Virtual Character Animation
- Multi-task Learning with Gradient Guided Policy Specialization
- Merging Deterministic Policy Gradient Estimations with Varied Bias-Variance Tradeoff for Effective Deep Reinforcement Learning
- Fractional Transfer Learning for Deep Model-Based Reinforcement Learning
- Distill Knowledge in Multi-task Reinforcement Learning with Optimal-Transport Regularization
- Interleaved Multitask Learning with Energy Modulated Learning Progress
- Efficient Reinforcement Learning in Resource Allocation Problems Through Permutation Invariant Multi-task Learning
- Learning Shared Dynamics with Meta-World Models