2 papers
cs.LG2024
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators
Allen Nie, Yash Chandak, Christina J. Yuan +3
Offline policy evaluation (OPE) allows us to evaluate and estimate a new sequential decision-making policy's performance by leveraging historical interaction data collected from ot…
cs.LG2021
SOPE: Spectrum of Off-Policy Estimators
Christina J. Yuan, Yash Chandak, Stephen Giguere +2
Many sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the…