Thompson Sampling is Asymptotically Optimal in General Environments
arXiv:1602.07905
Abstract
We discuss a variant of Thompson sampling for nonparametric reinforcement learning in a countable classes of general stochastic environments. These environments can be non-Markov, non-ergodic, and partially observable. We show that Thompson sampling learns the environment class in the sense that (1) asymptotically its value converges to the optimal value in mean and (2) given a recoverability assumption regret is sublinear.
UAI 2016
Cited by in corpus (8)
- Unifying Count-Based Exploration and Intrinsic Motivation
- Taming Non-stationary Bandits: A Bayesian Approach
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration
- Resolving Spurious Correlations in Causal Models of Environments via Interventions
- A Formal Solution to the Grain of Truth Problem
- Self-Modification of Policy and Utility Function in Rational Agents
- Value Directed Exploration in Multi-Armed Bandits with Structured Priors
- Curiosity Killed or Incapacitated the Cat and the Asymptotically Optimal Agent