Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
arXiv:1511.03722
Abstract
We study the problem of off-policy value evaluation in reinforcement learning (RL), where one aims to estimate the value of a new policy based on data collected by a different policy. This problem is often a critical step when applying RL in real-world problems. Despite its importance, existing general methods either have uncontrolled bias or suffer high variance. In this work, we extend the doubly robust estimator for bandits to sequential decision-making problems, which gets the best of both worlds: it is guaranteed to be unbiased and can have a much lower variance than the popular importance sampling estimators. We demonstrate the estimator's accuracy in several benchmark problems, and illustrate its use as a subroutine in safe policy improvement. We also provide theoretical results on the hardness of the problem, and show that our estimator can match the lower bound in certain scenarios.
14 pages; 4 figures; ICML 2016
References in corpus (3)
Cited by in corpus (25)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- A Survey on AI Assurance
- Constrained Policy Optimization
- Continuous State-Space Models for Optimal Sepsis Treatment - a Deep Reinforcement Learning Approach
- Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
- Representation and Reinforcement Learning for Personalized Glycemic Control in Septic Patients
- Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
- A Bayesian Nonparametric Approach for Estimating Individualized Treatment-Response Curves
- Off-Policy Evaluation via Off-Policy Classification
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical Activity
- Behaviour Policy Estimation in Off-Policy Policy Evaluation: Calibration Matters
- Off-Policy Evaluation via the Regularized Lagrangian
- Multi-step Off-policy Learning Without Importance Sampling Ratios
- Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning
- Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare? A Sensitivity Analysis of Duel-DDQN for Hemodynamic Management in Sepsis Patients
- On Ensuring that Intelligent Machines Are Well-Behaved
- Off-Policy Estimation of Long-Term Average Outcomes with Applications to Mobile Health
- Offline Policy Selection under Uncertainty
- Task Selection Policies for Multitask Learning
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs
- Imitation-Regularized Offline Learning
- Latent-state models for precision medicine
- Missingness as Stability: Understanding the Structure of Missingness in Longitudinal EHR data and its Impact on Reinforcement Learning in Healthcare
- Dynamic Measurement Scheduling for Adverse Event Forecasting using Deep RL
- Understanding the Artificial Intelligence Clinician and optimal treatment strategies for sepsis in intensive care