1 paper
Sutanoy Dasgupta, Yabo Niu, Kishan Panaganti +3
We consider the off-policy evaluation (OPE) problem in contextual bandits, where the goal is to estimate the value of a target policy using the data collected by a logging policy.…