1 citations · 1 across the 1 of their papers we have counts for
3 papers · 1 filter
Offline Evaluation of Reward-Optimizing Recommender Systems: The Case of Simulation
Imad Aouali, Amine Benhalloum, Martin Bompaire +5
Both in academic and industry-based research, online evaluation methods are seen as the golden standard for interactive applications like recommendation systems. Naturally, the rea…
Learning from Bandit Feedback: An Overview of the State-of-the-art
Olivier Jeunen, Dmytro Mykhaylov, David Rohde +3
In machine learning we often try to optimise a decision rule that would have worked well over a historical dataset; this is the so called empirical risk minimisation principle. In…
Three Methods for Training on Bandit Feedback
Dmytro Mykhaylov, David Rohde, Flavian Vasile +2
There are three quite distinct ways to train a machine learning model on recommender system logs. The first method is to model the reward prediction for each possible recommendatio…