23 citations · 37 across the 9 of their papers we have counts for
3 papers · 1 filter
Problem Dependent View on Structured Thresholding Bandit Problems
James Cheshire, Pierre Ménard, Alexandra Carpentier
We investigate the problem dependent regime in the stochastic Thresholding Bandit problem (TBP) under several shape constraints. In the TBP, the objective of the learner is to outp…
Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall
Tadashi Kozuno, Pierre Ménard, Rémi Munos +1
We study the problem of learning a Nash equilibrium (NE) in an imperfect information game (IIG) through self-play. Precisely, we focus on two-player, zero-sum, episodic, tabular II…
Gamification of Pure Exploration for Linear Bandits
Rémy Degenne, Pierre Ménard, Xuedong Shang +1
We investigate an active pure-exploration setting, that includes best-arm identification, in the context of linear stochastic bandits. While asymptotically optimal algorithms exist…