Policy Distillation
arXiv:1511.06295
Abstract
Policies for complex visual tasks have been successfully learned with deep reinforcement learning, using an approach called deep Q-networks (DQN), but relatively large (task-specific) networks and extensive training are needed to achieve good performance. In this work, we present a novel method called policy distillation that can be used to extract the policy of a reinforcement learning agent and train a new network that performs at the expert level while being dramatically smaller and more efficient. Furthermore, the same method can be used to consolidate multiple task-specific policies into a single policy. We demonstrate these claims using the Atari domain and show that the multi-task distilled agent outperforms the single-task teachers as well as a jointly-trained DQN agent.
Submitted to ICLR 2016
Cited by in corpus (52)
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- A Survey on Multi-Task Learning
- Born Again Neural Networks
- Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems
- Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
- Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
- Gradient Surgery for Multi-Task Learning
- Pseudo-Rehearsal: Achieving Deep Reinforcement Learning without Catastrophic Forgetting
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
- Robust Reinforcement Learning for Continuous Control with Model Misspecification
- Deep Reinforcement Learning with Successor Features for Navigation across Similar Environments
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- MoËT: Mixture of Expert Trees and its Application to Verifiable Reinforcement Learning
- Reinforcement Learning for Mixed Autonomy Intersections
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous Control
- Multi-Task Reinforcement Learning with Context-based Representations
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- Behavior Self-Organization Supports Task Inference for Continual Robot Learning
- Why distillation helps: a statistical perspective
- Self-Attentional Credit Assignment for Transfer in Reinforcement Learning
- Deep Learning for Embodied Vision Navigation: A Survey
- Online Robustness Training for Deep Reinforcement Learning
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- Simplifying Deep Reinforcement Learning via Self-Supervision
- Dual Policy Distillation
- From Few to More: Large-scale Dynamic Multiagent Curriculum Learning
- ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations
- Lifelong Robotic Reinforcement Learning by Retaining Experiences
- Actor Critic with Differentially Private Critic
- Techniques for Symbol Grounding with SATNet
- Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation
- Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation
- Optimal Completion Distillation for Sequence Learning
- Efficient Deep Reinforcement Learning via Adaptive Policy Transfer
- Preventing Posterior Collapse with Levenshtein Variational Autoencoder
- AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NER
- Synthesising Reinforcement Learning Policies through Set-Valued Inductive Rule Learning
- My Body is a Cage: the Role of Morphology in Graph-Based Incompatible Control
- Meta-CoTGAN: A Meta Cooperative Training Paradigm for Improving Adversarial Text Generation
- Developing Multi-Task Recommendations with Long-Term Rewards via Policy Distilled Reinforcement Learning
- Leveraging Undiagnosed Data for Glaucoma Classification with Teacher-Student Learning
- An Active Learning Framework for Efficient Robust Policy Search
- Learning in the Machine: the Symmetries of the Deep Learning Channel
- Training Over-parameterized Models with Non-decomposable Objectives
- Shared Learning : Enhancing Reinforcement in -Ensembles
- DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning
- On Estimating the Training Cost of Conversational Recommendation Systems
- When in Doubt, Summon the Titans: Efficient Inference with Large Models
- Interpretable Few-Shot Learning via Linear Distillation
- Transfer Value Iteration Networks