Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs
arXiv:2010.00587
Abstract
We study the reinforcement learning problem for discounted Markov Decision Processes (MDPs) under the tabular setting. We propose a model-based algorithm named UCBVI-, which is based on the \emph{optimism in the face of uncertainty principle} and the Bernstein-type bonus. We show that UCBVI- achieves an regret, where is the number of states, is the number of actions, is the discount factor and is the number of steps. In addition, we construct a class of hard MDPs and show that for any algorithm, the expected regret is at least . Our upper bound matches the minimax lower bound up to logarithmic factors, which suggests that UCBVI- is nearly minimax optimal for discounted MDPs.
33 pages, 1 figure, 1 table. In NeurIPS 2021
References in corpus (9)
- Empirical Bernstein Bounds and Sample Variance Penalization
- Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds
- On Lower Bounds for Regret in Reinforcement Learning
- Almost Optimal Model-Free Reinforcement Learning via Reference-Advantage Decomposition
- Variance-reduced -learning is minimax optimal
- Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP
- Model-Based Reinforcement Learning with a Generative Model is Minimax Optimal
- Model-Free Reinforcement Learning: from Clipped Pseudo-Regret to Sample Complexity
- Regret Bounds for Discounted MDPs