4 papers
EVaR-Optimal Arm Identification in Bandits
Mehrasa Ahmadipour, Aurélien Garivier
We study the fixed-confidence best arm identification (BAI) problem within the multi-armed bandit (MAB) framework under the Entropic Value-at-Risk (EVaR) criterion. Our analysis co…
Efficient Risk-sensitive Planning via Entropic Risk Measures
Alexandre Marthe, Samuel Bounan, Aurélien Garivier +1
Risk-sensitive planning aims to identify policies maximizing some tail-focused metrics in Markov Decision Processes (MDPs). Such an optimization task can be very costly for the mos…
Identifying the Best Transition Law
Mehrasa Ahmadipour, élise Crepon, Aurélien Garivier
Motivated by recursive learning in Markov Decision Processes, this paper studies best-arm identification in bandit problems where each arm's reward is drawn from a multinomial dist…
Sequential Learning of the Pareto Front for Multi-objective Bandits
Elise Crépon, Aurélien Garivier, Wouter M Koolen
We study the problem of sequential learning of the Pareto front in multi-objective multi-armed bandits. An agent is faced with K possible arms to pull. At each turn she picks one,…