2 papers
cs.IR2019
Learning from Bandit Feedback: An Overview of the State-of-the-art
Olivier Jeunen, Dmytro Mykhaylov, David Rohde +3
In machine learning we often try to optimise a decision rule that would have worked well over a historical dataset; this is the so called empirical risk minimisation principle. In…
cs.IR2019
Three Methods for Training on Bandit Feedback
Dmytro Mykhaylov, David Rohde, Flavian Vasile +2
There are three quite distinct ways to train a machine learning model on recommender system logs. The first method is to model the reward prediction for each possible recommendatio…