Deep Exploration via Randomized Value Functions
arXiv:1703.07608
Abstract
We study the use of randomized value functions to guide deep exploration in reinforcement learning. This offers an elegant means for synthesizing statistically and computationally efficient exploration with common practical approaches to value function learning. We present several reinforcement learning algorithms that leverage randomized value functions and demonstrate their efficacy through computational studies. We also prove a regret bound that establishes statistical efficiency with a tabular representation.
Accepted for publication in Journal of Machine Learning Research 2019
Cited by in corpus (59)
- Noisy Networks for Exploration
- Exploration in Deep Reinforcement Learning: A Survey
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- The Uncertainty Bellman Equation and Exploration
- Uncertainty in Neural Networks: Approximately Bayesian Ensembling
- ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation
- A Tutorial on Thompson Sampling
- Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder Dimension
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Randomized Value Functions via Multiplicative Normalizing Flows
- Worst-Case Regret Bounds for Exploration via Randomized Value Functions
- Variational Bayesian Reinforcement Learning with Regret Bounds
- Nonstationary Reinforcement Learning with Linear Function Approximation
- Revisiting Design Choices in Proximal Policy Optimization
- Statistical Bootstrapping for Uncertainty Estimation in Off-Policy Evaluation
- Optimistic Exploration even with a Pessimistic Initialisation
- Making Sense of Reinforcement Learning and Probabilistic Inference
- Provably Efficient Reinforcement Learning with Aggregated States
- MADE: Exploration via Maximizing Deviation from Explored Regions
- Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences
- Exploration by Distributional Reinforcement Learning
- Learning Adaptive Exploration Strategies in Dynamic Environments Through Informed Policy Regularization
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Adaptive Approximate Policy Iteration
- Offline Policy Selection under Uncertainty
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Effective Exploration for Deep Reinforcement Learning via Bootstrapped Q-Ensembles under Tsallis Entropy Regularization
- Temporal Difference Uncertainties as a Signal for Exploration
- Principled Exploration via Optimistic Bootstrapping and Backward Induction
- A Survey of Exploration Methods in Reinforcement Learning
- Reinforcement Learning, Bit by Bit
- Discovery of Options via Meta-Learned Subgoals
- Targeting for long-term outcomes
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated Environments
- Regret Bounds for Stochastic Shortest Path Problems with Linear Function Approximation
- Bayesian Neural Network Ensembles
- Learning more skills through optimistic exploration
- Improving the Diversity of Bootstrapped DQN by Replacing Priors With Noise
- Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
- Randomized Value Functions via Posterior State-Abstraction Sampling
- An Information-Theoretic Perspective on Credit Assignment in Reinforcement Learning
- Efficient exploration of zero-sum stochastic games
- Learning Diverse Policies with Soft Self-Generated Guidance
- Learning to Be Cautious
- Deciding What to Learn: A Rate-Distortion Approach
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- Dr Jekyll and Mr Hyde: the Strange Case of Off-Policy Policy Updates
- TTR-Based Reward for Reinforcement Learning with Implicit Model Priors
- Gaussian-Dirichlet Posterior Dominance in Sequential Learning
- Improved Algorithms for Misspecified Linear Markov Decision Processes
- Continuous Control With Ensemble Deep Deterministic Policy Gradients
- Optimal Network Control in Partially-Controllable Networks
- Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning
- Deep Exploration for Recommendation Systems
- Bayesian Distributional Policy Gradients
- The Value of Information When Deciding What to Learn
- Amortized Variational Deep Q Network
- Distributed interference cancellation in multi-agent scenarios