2 papers
cs.LG2026
Top- Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection
Nicolas Gutowski, Fabien Chhel, Alexandre Letard +1
We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of arms and observes their -dimensional reward vectors under semi-ban…
cs.LG2020
Partial Bandit and Semi-Bandit: Making the Most Out of Scarce Users' Feedback
Alexandre Letard, Tassadit Amghar, Olivier Camp +1
Recent works on Multi-Armed Bandits (MAB) and Combinatorial Multi-Armed Bandits (COM-MAB) show good results on a global accuracy metric. This can be achieved, in the case of recomm…