Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
arXiv:2108.08812
Abstract
Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that naturally incorporates the pessimism principle, leading to several key advantages compared to the state of the art. The algorithm can operate when the Bellman evaluation operator is closed with respect to the action value function of the actor's policies; this is a more general setting than the low-rank MDP model. Despite the added generality, the procedure is computationally tractable as it involves the solution of a sequence of second-order programs. We prove an upper bound on the suboptimality gap of the policy returned by the procedure that depends on the data coverage of any arbitrary, possibly data dependent comparator policy. The achievable guarantee is complemented with a minimax lower bound that is matching up to logarithmic factors.
Initial submission; appeared as spotlight talk in ICML 2021 Workshop on Theory of RL
References in corpus (14)
- Conservative Q-Learning for Offline Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- Contextual Decision Processes with Low Bellman Rank are PAC-Learnable
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- A Theory of Regularized Markov Decision Processes
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Variational Policy Gradient Method for Reinforcement Learning with General Utilities
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPs
- Provably Good Batch Reinforcement Learning Without Great Exploration
- Reinforcement Learning via Fenchel-Rockafellar Duality
- Off-Policy Evaluation via the Regularized Lagrangian
- Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
- Minimax-Optimal Off-Policy Evaluation with Linear Function Approximation
- Risk Bounds and Rademacher Complexity in Batch Reinforcement Learning