Posterior sampling for reinforcement learning: worst-case regret bounds
arXiv:1705.07041
Abstract
We present an algorithm based on posterior sampling (aka Thompson sampling) that achieves near-optimal worst-case regret bounds when the underlying Markov Decision Process (MDP) is communicating with a finite, though unknown, diameter. Our main result is a high probability regret upper bound of for any communicating MDP with states, actions and diameter . Here, regret compares the total reward achieved by the algorithm to the total expected reward of an optimal infinite-horizon undiscounted average reward policy, in time horizon . This result closely matches the known lower bound of . Our techniques involve proving some novel results about the anti-concentration of Dirichlet distribution, which may be of independent interest.
This revision fixes an error due to use of some incorrect results (Lemma C.1 and Lemma C.2) in the earlier version. The regret bounds in this version are worse by a factor of sqrt(S) as compared to the previous version
References in corpus (6)
- Further Optimal Regret Bounds for Thompson Sampling
- (More) Efficient Reinforcement Learning via Posterior Sampling
- A Bayesian Sampling Approach to Exploration in Reinforcement Learning
- REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs
- Sample Complexity of Episodic Fixed-Horizon Reinforcement Learning
- Variance Reduction Methods for Sublinear Reinforcement Learning