activity
20192022
most citedGamification of Pure Exploration for Linear Bandits

23 citations · 37 across the 8 of their papers we have counts for

collaborators

12 papers

cs.LG2022

KL-Entropy-Regularized RL with a Generative Model is Minimax Optimal

Tadashi Kozuno, Wenhao Yang, Nino Vieillard +10

In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model. Particularly, we analyze mirror descent value iteration (M…

stat.ML20211 cited

Problem Dependent View on Structured Thresholding Bandit Problems

James Cheshire, Pierre Ménard, Alexandra Carpentier

We investigate the problem dependent regime in the stochastic Thresholding Bandit problem (TBP) under several shape constraints. In the TBP, the objective of the learner is to outp…

stat.ML20212 cited

Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall

Tadashi Kozuno, Pierre Ménard, Rémi Munos +1

We study the problem of learning a Nash equilibrium (NE) in an imperfect information game (IIG) through self-play. Precisely, we focus on two-player, zero-sum, episodic, tabular II…

cs.LG2021

Bandits with many optimal arms

Rianne de Heide, James Cheshire, Pierre Ménard +1

We consider a stochastic bandit problem with a possibly infinite number of arms. We write for the proportion of optimal arms and for the minimal mean-gap between optimal…

cs.LG2020

Episodic Reinforcement Learning in Finite MDPs: Minimax Lower Bounds Revisited

Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann +1

In this paper, we propose new problem-independent lower bounds on the sample complexity and regret in episodic MDPs, with a particular focus on the non-stationary case in which the…

cs.IT2020

Optimal Strategies for Graph-Structured Bandits

Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard

We study a structured variant of the multi-armed bandit problem specified by a set of Bernoulli distributions with mean…