23 citations · 38 across the 14 of their papers we have counts for
5 papers · 1 filter
Indexed Minimum Empirical Divergence for Unimodal Bandits
Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard
We consider a multi-armed bandit problem specified by a set of one-dimensional family exponential distributions endowed with a unimodal structure. We introduce IMED-UB, a algorithm…
Adaptive Multi-Goal Exploration
Jean Tarbouriech, Omar Darwiche Domingues, Pierre Ménard +3
We introduce a generic strategy for provably efficient multi-goal exploration. It relies on AdaGoal, a novel goal selection scheme that leverages a measure of uncertainty in reachi…
Problem Dependent View on Structured Thresholding Bandit Problems
James Cheshire, Pierre Ménard, Alexandra Carpentier
We investigate the problem dependent regime in the stochastic Thresholding Bandit problem (TBP) under several shape constraints. In the TBP, the objective of the learner is to outp…
Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall
Tadashi Kozuno, Pierre Ménard, Rémi Munos +1
We study the problem of learning a Nash equilibrium (NE) in an imperfect information game (IIG) through self-play. Precisely, we focus on two-player, zero-sum, episodic, tabular II…
Bandits with many optimal arms
Rianne de Heide, James Cheshire, Pierre Ménard +1
We consider a stochastic bandit problem with a possibly infinite number of arms. We write for the proportion of optimal arms and for the minimal mean-gap between optimal…