2 papers
cs.LG2025
Identifying the Best Transition Law
Mehrasa Ahmadipour, élise Crepon, Aurélien Garivier
Motivated by recursive learning in Markov Decision Processes, this paper studies best-arm identification in bandit problems where each arm's reward is drawn from a multinomial dist…
stat.ML2025
Sequential Learning of the Pareto Front for Multi-objective Bandits
Elise Crépon, Aurélien Garivier, Wouter M Koolen
We study the problem of sequential learning of the Pareto front in multi-objective multi-armed bandits. An agent is faced with K possible arms to pull. At each turn she picks one,…