1 paper
Yuta Saito, Jihan Yao, Thorsten Joachims
We study off-policy learning (OPL) of contextual bandit policies in large discrete action spaces where existing methods -- most of which rely crucially on reward-regression models…