1 paper
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
Off-policy learning (OPL) aims at finding improved policies from logged bandit data, often by minimizing the inverse propensity scoring (IPS) estimator of the risk. In this work, w…