3 papers
cs.IR2019
Learning from Bandit Feedback: An Overview of the State-of-the-art
Olivier Jeunen, Dmytro Mykhaylov, David Rohde +3
In machine learning we often try to optimise a decision rule that would have worked well over a historical dataset; this is the so called empirical risk minimisation principle. In…
cs.IR2019
Three Methods for Training on Bandit Feedback
Dmytro Mykhaylov, David Rohde, Flavian Vasile +2
There are three quite distinct ways to train a machine learning model on recommender system logs. The first method is to model the reward prediction for each possible recommendatio…
stat.ML2018
Dual optimization for convex constrained objectives without the gradient-Lipschitz assumption
Martin Bompaire, Emmanuel Bacry, Stéphane Gaïffas
The minimization of convex objectives coming from linear supervised learning problems, such as penalized generalized linear models, can be formulated as finite sums of convex funct…