45 citations · 47 across the 3 of their papers we have counts for
3 papers · 2 filters
Behaviour Policy Estimation in Off-Policy Policy Evaluation: Calibration Matters
Aniruddh Raghu, Omer Gottesman, Yao Liu +4
In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empi…
Evaluating Reinforcement Learning Algorithms in Observational Health Settings
Omer Gottesman, Fredrik Johansson, Joshua Meier +16
Much attention has been devoted recently to the development of machine learning algorithms with the goal of improving treatment policies in healthcare. Reinforcement learning (RL)…
Representation Balancing MDPs for Off-Policy Policy Evaluation
Yao Liu, Omer Gottesman, Aniruddh Raghu +4
We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast to prior work, we consider how to estimate both the individual policy value and average policy value…