1 citations · 1 across the 1 of their papers we have counts for
1 paper
Ben London, Levi Lu, Ted Sandler +1
We propose the first boosting algorithm for off-policy learning from logged bandit feedback. Unlike existing boosting methods for supervised learning, our algorithm directly optimi…