9 citations · 16 across the 4 of their papers we have counts for
3 papers · 1 filter
Deeply-Debiased Off-Policy Interval Estimation
Chengchun Shi, Runzhe Wan, Victor Chernozhukov +1
Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy. In addition to a point estimate, many applications would be…
Deep Jump Learning for Off-Policy Evaluation in Continuous Treatment Settings
Hengrui Cai, Chengchun Shi, Rui Song +1
We consider off-policy evaluation (OPE) in continuous treatment settings, such as personalized dose-finding. In OPE, one aims to estimate the mean outcome under a new treatment dec…
Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision Making
Chengchun Shi, Runzhe Wan, Rui Song +2
The Markov assumption (MA) is fundamental to the empirical validity of reinforcement learning. In this paper, we propose a novel Forward-Backward Learning procedure to test MA in s…