Bootstrapped Thompson Sampling and Deep Exploration
arXiv:1507.00300
Abstract
This technical note presents a new approach to carrying out the kind of exploration achieved by Thompson sampling, but without explicitly maintaining or sampling from posterior distributions. The approach is based on a bootstrap technique that uses a combination of observed and artificially generated data. The latter serves to induce a prior distribution which, as we will demonstrate, is critical to effective exploration. We explain how the approach can be applied to multi-armed bandit and reinforcement learning problems and how it relates to Thompson sampling. The approach is particularly well-suited for contexts in which exploration is coupled with deep learning, since in these settings, maintaining or generating samples from a posterior distribution becomes computationally infeasible.
References in corpus (3)
Cited by in corpus (7)
- Variational Deep Q Network
- Risk Aware and Multi-Objective Decision Making with Distributional Monte Carlo Tree Search
- Residual Bootstrap Exploration for Bandit Algorithms
- Differentiable Linear Bandit Algorithm
- CORe: Capitalizing On Rewards in Bandit Exploration
- State-Aware Variational Thompson Sampling for Deep Q-Networks
- Knowledge is reward: Learning optimal exploration by predictive reward cashing