Ensemble Sampling
arXiv:1705.07347
Abstract
Thompson sampling has emerged as an effective heuristic for a broad range of online decision problems. In its basic form, the algorithm requires computing and sampling from a posterior distribution over models, which is tractable only for simple special cases. This paper develops ensemble sampling, which aims to approximate Thompson sampling while maintaining tractability even in the face of complex models such as neural networks. Ensemble sampling dramatically expands on the range of applications for which Thompson sampling is viable. We establish a theoretical basis that supports the approach and present computational results that offer further insight.
Cited by in corpus (21)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Uncertainty in Neural Networks: Approximately Bayesian Ensembling
- A Tutorial on Thompson Sampling
- Behaviour Suite for Reinforcement Learning
- Garbage In, Reward Out: Bootstrapping Exploration in Multi-Armed Bandits
- Randomized Exploration in Generalized Linear Bandits
- Hypermodels for Exploration
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Perturbed-History Exploration in Stochastic Linear Bandits
- Empirical Bayes Regret Minimization
- Targeting for long-term outcomes
- Bayesian Neural Network Ensembles
- On Thompson Sampling for Smoother-than-Lipschitz Bandits
- Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
- Randomized Value Functions via Posterior State-Abstraction Sampling
- Coordinated Exploration in Concurrent Reinforcement Learning
- Neural Model-based Optimization with Right-Censored Observations
- Analysis and Design of Thompson Sampling for Stochastic Partial Monitoring
- Learning to Be Cautious
- Deep Exploration for Recommendation Systems
- The Value of Information When Deciding What to Learn