Beyond No Regret: Instance-Dependent PAC Reinforcement Learning
arXiv:2108.02717
Abstract
The theory of reinforcement learning has focused on two fundamental problems: achieving low regret, and identifying -optimal policies. While a simple reduction allows one to apply a low-regret algorithm to obtain an -optimal policy and achieve the worst-case optimal rate, it is unknown whether low-regret algorithms can obtain the instance-optimal rate for policy identification. We show this is not possible -- there exists a fundamental tradeoff between achieving low regret and identifying an -optimal policy at the instance-optimal rate. Motivated by our negative finding, we propose a new measure of instance-dependent sample complexity for PAC tabular reinforcement learning which explicitly accounts for the attainable state visitation distributions in the underlying MDP. We then propose and analyze a novel, planning-based algorithm which attains this sample complexity -- yielding a complexity which scales with the suboptimality gaps and the "reachability" of a state. We show our algorithm is nearly minimax optimal, and on several examples that our instance-dependent sample complexity offers significant improvements over worst-case bounds.
References in corpus (12)
- Empirical Bernstein Bounds and Sample Variance Penalization
- Fast active learning for pure exploration in reinforcement learning
- Is Reinforcement Learning More Difficult Than Bandits? A Near-optimal Algorithm Escaping the Curse of Horizon
- Reward-Free Exploration for Reinforcement Learning
- Is Long Horizon Reinforcement Learning More Difficult Than Short Horizon Reinforcement Learning?
- Fine-Grained Gap-Dependent Bounds for Tabular MDPs via Adaptive Multi-Step Bootstrap
- Nearly Minimax Optimal Reward-free Reinforcement Learning
- Adaptive Sampling for Best Policy Identification in Markov Decision Processes
- Navigating to the Best Policy in Markov Decision Processes
- Planning in Markov Decision Processes with Gap-Dependent Sample Complexity
- Task-Optimal Exploration in Linear Dynamical Systems
- Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning