Deep Q-learning from Demonstrations
arXiv:1704.03732
Abstract
Deep reinforcement learning (RL) has achieved several high profile successes in difficult decision-making problems. However, these algorithms typically require a huge amount of data before they reach reasonable performance. In fact, their performance during learning can be extremely poor. This may be acceptable for a simulator, but it severely limits the applicability of deep RL to many real-world tasks, where the agent must learn in the real environment. In this paper we study a setting where the agent may access data from previous control of the system. We present an algorithm, Deep Q-learning from Demonstrations (DQfD), that leverages small sets of demonstration data to massively accelerate the learning process even from relatively small amounts of demonstration data and is able to automatically assess the necessary ratio of demonstration data while learning thanks to a prioritized replay mechanism. DQfD works by combining temporal difference updates with supervised classification of the demonstrator's actions. We show that DQfD has better initial performance than Prioritized Dueling Double Deep Q-Networks (PDD DQN) as it starts with better scores on the first million steps on 41 of 42 games and on average it takes PDD DQN 83 million steps to catch up to DQfD's performance. DQfD learns to out-perform the best demonstration given in 14 of 42 games. In addition, DQfD leverages human demonstrations to achieve state-of-the-art results for 11 games. Finally, we show that DQfD performs better than three related algorithms for incorporating demonstration data into DQN.
Published at AAAI 2018. Previously on arxiv as "Learning from Demonstrations for Real World Reinforcement Learning"
Cited by in corpus (135)
- Informed Machine Learning -- A Taxonomy and Survey of Integrating Knowledge into Learning Systems
- StarCraft II: A New Challenge for Reinforcement Learning
- Conservative Q-Learning for Offline Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- Neo: A Learned Query Optimizer
- Experience Replay for Continual Learning
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Off-Policy Deep Reinforcement Learning without Exploration
- Challenges of Real-World Reinforcement Learning
- Go-Explore: a New Approach for Hard-Exploration Problems
- PRIMAL2: Pathfinding via Reinforcement and Imitation Multi-Agent Learning -- Lifelong
- Transfer Learning in Deep Reinforcement Learning: A Survey
- A survey on intrinsic motivation in reinforcement learning
- Observe and Look Further: Achieving Consistent Performance on Atari
- One for Many: Transfer Learning for Building HVAC Control
- Deep reinforcement learning for guidewire navigation in coronary artery phantom
- Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
- The MineRL 2019 Competition on Sample Efficient Reinforcement Learning using Human Priors
- Kickstarting Deep Reinforcement Learning
- Primal Wasserstein Imitation Learning
- Transfer Learning for Future Wireless Networks: A Comprehensive Survey
- Reinforcement Learning for Robotic Manipulation using Simulated Locomotion Demonstrations
- Learning Variable Ordering Heuristics for Solving Constraint Satisfaction Problems
- Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations
- BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning
- Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
- Human-in-the-Loop Deep Reinforcement Learning with Application to Autonomous Driving
- Learning to Reach Goals via Iterated Supervised Learning
- Hierarchical Imitation and Reinforcement Learning
- Backplay: "Man muss immer umkehren"
- Penalizing side effects using stepwise relative reachability
- Learning latent state representation for speeding up exploration
- Constrained-Space Optimization and Reinforcement Learning for Complex Tasks
- What Matters for Adversarial Imitation Learning?
- A View on Deep Reinforcement Learning in System Optimization
- Distributionally Robust Reinforcement Learning
- Toward the Fundamental Limits of Imitation Learning
- State Alignment-based Imitation Learning
- AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- Safety-Guided Deep Reinforcement Learning via Online Gaussian Process Estimation
- CopyCAT: Taking Control of Neural Policies with Constant Attacks
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
- Watch, Try, Learn: Meta-Learning from Demonstrations and Reward
- Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset
- Heuristic-Guided Reinforcement Learning
- Flatland-RL : Multi-Agent Reinforcement Learning on Trains
- Batch Policy Learning under Constraints
- Student-Initiated Action Advising via Advice Novelty
- Efficient Deep Reinforcement Learning with Imitative Expert Priors for Autonomous Driving
- Residual Reinforcement Learning from Demonstrations
- Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
- Dual Policy Distillation
- Imitation Learning: Progress, Taxonomies and Challenges
- Security Issues of Low Power Wide Area Networks in the Context of LoRa Networks
- Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments
- Adversarial Exploitation of Policy Imitation
- Reinforcement Learning from Imperfect Demonstrations under Soft Expert Guidance
- Demonstration-Guided Reinforcement Learning with Learned Skills
- Sample Efficient Reinforcement Learning through Learning from Demonstrations in Minecraft
- Grasping in the Wild:Learning 6DoF Closed-Loop Grasping from Low-Cost Demonstrations
- ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations
- Modeling 3D Shapes by Reinforcement Learning
- From Few to More: Large-scale Dynamic Multiagent Curriculum Learning
- Safety Augmented Value Estimation from Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
- Scaling Imitation Learning in Minecraft
- Reinforcement Learning with Supervision from Noisy Demonstrations
- A Deep Q-learning/genetic Algorithms Based Novel Methodology For Optimizing Covid-19 Pandemic Government Actions
- Learning hierarchical behavior and motion planning for autonomous driving
- I'm sorry Dave, I'm afraid I can't do that, Deep Q-learning from forbidden action
- Shaping Rewards for Reinforcement Learning with Imperfect Demonstrations using Generative Models
- Human AI interaction loop training: New approach for interactive reinforcement learning
- Hindsight Generative Adversarial Imitation Learning
- Exploration-efficient Deep Reinforcement Learning with Demonstration Guidance for Robot Control
- Policy Learning Using Weak Supervision
- PsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference Learning
- Pre-training in Deep Reinforcement Learning for Automatic Speech Recognition
- Equivariant Learning in Spatial Action Spaces
- Robust Multi-Modal Policies for Industrial Assembly via Reinforcement Learning and Demonstrations: A Large-Scale Study
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Policies
- Stealing Deep Reinforcement Learning Models for Fun and Profit
- The Chef's Hat Simulation Environment for Reinforcement-Learning-Based Agents
- Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
- Playing Minecraft with Behavioural Cloning
- Imitation Learning from Pixel-Level Demonstrations by HashReward
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- Variance-Reduced Off-Policy Memory-Efficient Policy Search
- Reinforced Imitation Learning by Free Energy Principle
- Co-training for Policy Learning
- Distributed Resource Scheduling for Large-Scale MEC Systems: A Multi-Agent Ensemble Deep Reinforcement Learning with Imitation Acceleration
- Bridging the Imitation Gap by Adaptive Insubordination
- Policy Gradient from Demonstration and Curiosity
- Follow the Object: Curriculum Learning for Manipulation Tasks with Imagined Goals
- Co-Imitation Learning without Expert Demonstration
- Forgetful Experience Replay in Hierarchical Reinforcement Learning from Demonstrations
- Constrained Exploration and Recovery from Experience Shaping
- Fast Reinforcement Learning for Anti-jamming Communications
- Reinforcement Learning for Industrial Control Network Cyber Security Orchestration
- Guided Exploration with Proximal Policy Optimization using a Single Demonstration
- Continuous Control with Action Quantization from Demonstrations
- Plan-Based Relaxed Reward Shaping for Goal-Directed Tasks
- Reparameterized Variational Divergence Minimization for Stable Imitation
- Discovering an Aid Policy to Minimize Student Evasion Using Offline Reinforcement Learning
- SEIHAI: A Sample-efficient Hierarchical AI for the MineRL Competition
- Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation
- Few-Shot Bayesian Imitation Learning with Logical Program Policies
- Policy learning in SE(3) action spaces
- SREC: Proactive Self-Remedy of Energy-Constrained UAV-Based Networks via Deep Reinforcement Learning
- Towards robust and domain agnostic reinforcement learning competitions
- Tolerance-Guided Policy Learning for Adaptable and Transferrable Delicate Industrial Insertion
- Episodic Self-Imitation Learning with Hindsight
- Periodic Intra-Ensemble Knowledge Distillation for Reinforcement Learning
- Self-Imitation Learning for Robot Tasks with Sparse and Delayed Rewards
- Chrome Dino Run using Reinforcement Learning
- Constrained Policy Improvement for Safe and Efficient Reinforcement Learning
- Potential-based Reward Shaping in Sokoban
- An advantage actor-critic algorithm for robotic motion planning in dense and dynamic scenarios
- Deep reinforcement learning for RAN optimization and control
- Multi-Preference Actor Critic
- DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning
- Active Hierarchical Imitation and Reinforcement Learning
- Learning Memory-Dependent Continuous Control from Demonstrations
- Reinforcement Learning via Reasoning from Demonstration
- Evolutionary Stochastic Policy Distillation
- Responsive Regulation of Dynamic UAV Communication Networks Based on Deep Reinforcement Learning
- Make Bipedal Robots Learn How to Imitate
- Integration of Imitation Learning using GAIL and Reinforcement Learning using Task-achievement Rewards via Probabilistic Graphical Model
- Mixing Human Demonstrations with Self-Exploration in Experience Replay for Deep Reinforcement Learning
- Ranking Policy Gradient
- Sample Efficient Imitation Learning via Reward Function Trained in Advance
- RLCache: Automated Cache Management Using Reinforcement Learning
- Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy
- Soft-Robust Algorithms for Batch Reinforcement Learning