Variance-Aware Off-Policy Evaluation with Linear Function Approximation
arXiv:2106.11960
Abstract
We study the off-policy evaluation (OPE) problem in reinforcement learning with linear function approximation, which aims to estimate the value function of a target policy based on the offline data collected by a behavior policy. We propose to incorporate the variance information of the value function to improve the sample efficiency of OPE. More specifically, for time-inhomogeneous episodic linear Markov decision processes (MDPs), we propose an algorithm, VA-OPE, which uses the estimated variance of the value function to reweight the Bellman residual in Fitted Q-Iteration. We show that our algorithm achieves a tighter error bound than the best-known result. We also provide a fine-grained characterization of the distribution shift between the behavior policy and the target policy. Extensive numerical experiments corroborate our theory.
59 pages, 4 figures. In NeurIPS 2021
References in corpus (12)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Doubly Robust Policy Evaluation and Learning
- MOReL : Model-Based Offline Reinforcement Learning
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Model-Based Reinforcement Learning with Value-Targeted Regression
- Learning Near Optimal Policies with Low Inherent Bellman Error
- GenDICE: Generalized Offline Estimation of Stationary Values
- Is Pessimism Provably Efficient for Offline RL?
- Near-Optimal Provable Uniform Convergence in Offline Policy Evaluation for Reinforcement Learning
- CoinDICE: Off-Policy Confidence Interval Estimation
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation
- Infinite-Horizon Offline Reinforcement Learning with Linear Function Approximation: Curse of Dimensionality and Algorithm