1 paper · 1 filter
Mohammad Mehrabi, Stefan Wager
Doubly robust methods hold considerable promise for off-policy evaluation in Markov decision processes (MDPs) under sequential ignorability: They have been shown to converge as $1/…