Generalization and Exploration via Randomized Value Functions
arXiv:1402.0635
Abstract
We propose randomized least-squares value iteration (RLSVI) -- a new reinforcement learning algorithm designed to explore and generalize efficiently via linearly parameterized value functions. We explain why versions of least-squares value iteration that use Boltzmann or epsilon-greedy exploration can be highly inefficient, and we present computational results that demonstrate dramatic efficiency gains enjoyed by RLSVI. Further, we establish an upper bound on the expected regret of RLSVI that demonstrates near-optimality in a tabula rasa learning context. More broadly, our results suggest that randomized value functions offer a promising approach to tackling a critical challenge in reinforcement learning: synthesizing efficient exploration and effective generalization.
arXiv admin note: text overlap with arXiv:1307.4847
References in corpus (5)
- Further Optimal Regret Bounds for Thompson Sampling
- Bootstrapped Thompson Sampling and Deep Exploration
- Online Regret Bounds for Undiscounted Continuous Reinforcement Learning
- Model-based Reinforcement Learning and the Eluder Dimension
- State of the Art Control of Atari Games Using Shallow Reinforcement Learning
Cited by in corpus (36)
- Noisy Networks for Exploration
- Exploration in Deep Reinforcement Learning: A Survey
- VIME: Variational Information Maximizing Exploration
- Deep Exploration via Bootstrapped DQN
- Stochastic Neural Networks for Hierarchical Reinforcement Learning
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Model-Based Reinforcement Learning with Value-Targeted Regression
- Bootstrapped Thompson Sampling and Deep Exploration
- The Uncertainty Bellman Equation and Exploration
- Rule-Based Reinforcement Learning for Efficient Robot Navigation with Space Reduction
- Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?
- Provable Self-Play Algorithms for Competitive Reinforcement Learning
- Model-based Reinforcement Learning and the Eluder Dimension
- From Language to Programs: Bridging Reinforcement Learning and Maximum Marginal Likelihood
- Thompson Sampling for Linear-Quadratic Control Problems
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient Learning
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes
- Efficient exploration with Double Uncertain Value Networks
- Near-Optimal Reinforcement Learning with Self-Play
- Variational Bayesian Reinforcement Learning with Regret Bounds
- Provably Efficient Causal Reinforcement Learning with Confounded Observational Data
- V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL
- Assumed Density Filtering Q-learning
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- The Potential of the Return Distribution for Exploration in RL
- Optimal Demand Response Using Device Based Reinforcement Learning
- Learning Zero-Sum Simultaneous-Move Markov Games Using Function Approximation and Correlated Equilibrium
- Improving the Diversity of Bootstrapped DQN by Replacing Priors With Noise
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy Optimization
- Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement Learning
- Gaussian-Dirichlet Posterior Dominance in Sequential Learning
- Intrinsic Exploration as Multi-Objective RL
- Angrier Birds: Bayesian reinforcement learning