Deep Exploration via Bootstrapped DQN
arXiv:1602.04621
Abstract
Efficient exploration in complex environments remains a major challenge for reinforcement learning. We propose bootstrapped DQN, a simple algorithm that explores in a computationally and statistically efficient manner through use of randomized value functions. Unlike dithering strategies such as epsilon-greedy exploration, bootstrapped DQN carries out temporally-extended (or deep) exploration; this can lead to exponentially faster learning. We demonstrate these benefits in complex stochastic MDPs and in the large-scale Arcade Learning Environment. Bootstrapped DQN substantially improves learning times and performance across most Atari games.
References in corpus (5)
Cited by in corpus (95)
- A Brief Survey of Deep Reinforcement Learning
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- Exploration in Deep Reinforcement Learning: A Survey
- VIME: Variational Information Maximizing Exploration
- Parameter Space Noise for Exploration
- Large-Scale Study of Curiosity-Driven Learning
- Hindsight Experience Replay
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- Exploration by Random Network Distillation
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance
- Multiplicative Normalizing Flows for Variational Bayesian Neural Networks
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Meta-Reinforcement Learning of Structured Exploration Strategies
- CEM-RL: Combining evolutionary and gradient-based methods for policy search
- Neural Episodic Control
- Deep Successor Reinforcement Learning
- Learning values across many orders of magnitude
- Playing Atari Games with Deep Reinforcement Learning and Human Checkpoint Replay
- Weakly Supervised-Based Oversampling for High Imbalance and High Dimensionality Data Classification
- Stein Variational Policy Gradient
- Ensemble Quantile Networks: Uncertainty-Aware Reinforcement Learning with Applications in Autonomous Driving
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- The Uncertainty Bellman Equation and Exploration
- One-Shot Reinforcement Learning for Robot Navigation with Interactive Replay
- Unsupervised Meta-Learning for Reinforcement Learning
- Counting to Explore and Generalize in Text-based Games
- Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening
- Scalable Uncertainty Quantification for Deep Operator Networks using Randomized Priors
- A Smooth Representation of Belief over SO(3) for Deep Rotation Learning with Uncertainty
- Scalable End-to-End Autonomous Vehicle Testing via Rare-event Simulation
- From Language to Programs: Bridging Reinforcement Learning and Maximum Marginal Likelihood
- Towards a Common Implementation of Reinforcement Learning for Multiple Robotic Tasks
- AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates
- Neural Thompson Sampling
- Thompson Sampling for Linear-Quadratic Control Problems
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Multi-task Deep Reinforcement Learning with PopArt
- Learning to Explore with Meta-Policy Gradient
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- Asynchronous Parallel Bayesian Optimisation via Thompson Sampling
- NPBDREG: Uncertainty Assessment in Diffeomorphic Brain MRI Registration using a Non-parametric Bayesian Deep-Learning Based Approach
- Efficient exploration with Double Uncertain Value Networks
- L2C2: Locally Lipschitz Continuous Constraint towards Stable and Smooth Reinforcement Learning
- Bayesian Policy Gradients via Alpha Divergence Dropout Inference
- Deep Network Uncertainty Maps for Indoor Navigation
- Multi-Advisor Reinforcement Learning
- FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback
- Exploration by Distributional Reinforcement Learning
- Discovering Blind Spots in Reinforcement Learning
- Deep Reinforcement Learning With Macro-Actions
- ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning
- Assumed Density Filtering Q-learning
- Uncertainty Estimates for Efficient Neural Network-based Dialogue Policy Optimisation
- Adaptive Online Planning for Continual Lifelong Learning
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Variational Deep Q Network
- When and How Mixup Improves Calibration
- Novelty Search in Representational Space for Sample Efficient Exploration
- The Potential of the Return Distribution for Exploration in RL
- Uncertainty-sensitive Learning and Planning with Ensembles
- Strength in Numbers: Trading-off Robustness and Computation via Adversarially-Trained Ensembles
- Offline reinforcement learning with uncertainty for treatment strategies in sepsis
- Deep Learning in Robotics: A Review of Recent Research
- Developing an OpenAI Gym-compatible framework and simulation environment for testing Deep Reinforcement Learning agents solving the Ambulance Location Problem
- Challenges of Context and Time in Reinforcement Learning: Introducing Space Fortress as a Benchmark
- Deep Reinforcement Learning with Feedback-based Exploration
- Distributional Actor-Critic Ensemble for Uncertainty-Aware Continuous Control
- Reward Bonuses with Gain Scheduling Inspired by Iterative Deepening Search
- Distilling Knowledge for Search-based Structured Prediction
- PBCS : Efficient Exploration and Exploitation Using a Synergy between Reinforcement Learning and Motion Planning
- Improving the Diversity of Bootstrapped DQN by Replacing Priors With Noise
- The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction
- Making Curiosity Explicit in Vision-based RL
- Investigating and Improving Latent Density Segmentation Models for Aleatoric Uncertainty Quantification in Medical Imaging
- Sparse Attention Guided Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning
- Uncertainty Quantification From Scaling Laws in Deep Neural Networks
- ACE: An Actor Ensemble Algorithm for Continuous Control with Tree Search
- Merging Deterministic Policy Gradient Estimations with Varied Bias-Variance Tradeoff for Effective Deep Reinforcement Learning
- Dual policy as self-model for planning
- Uncertainty Propagation in Node Classification
- Improving width-based planning with compact policies
- Generative Adversarial Exploration for Reinforcement Learning
- Shared Learning : Enhancing Reinforcement in -Ensembles
- Uncertainty-Aware Data Aggregation for Deep Imitation Learning
- Progressive extension of reinforcement learning action dimension for asymmetric assembly tasks
- A Comparison of the Delta Method and the Bootstrap in Deep Learning Classification
- Planning with Exploration: Addressing Dynamics Bottleneck in Model-based Reinforcement Learning
- Modelling resource allocation in uncertain system environment through deep reinforcement learning
- Depth and nonlinearity induce implicit exploration for RL