23 citations · 37 across the 8 of their papers we have counts for
12 papers
KL-Entropy-Regularized RL with a Generative Model is Minimax Optimal
Tadashi Kozuno, Wenhao Yang, Nino Vieillard +10
In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model. Particularly, we analyze mirror descent value iteration (M…
Problem Dependent View on Structured Thresholding Bandit Problems
James Cheshire, Pierre Ménard, Alexandra Carpentier
We investigate the problem dependent regime in the stochastic Thresholding Bandit problem (TBP) under several shape constraints. In the TBP, the objective of the learner is to outp…
Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall
Tadashi Kozuno, Pierre Ménard, Rémi Munos +1
We study the problem of learning a Nash equilibrium (NE) in an imperfect information game (IIG) through self-play. Precisely, we focus on two-player, zero-sum, episodic, tabular II…
Bandits with many optimal arms
Rianne de Heide, James Cheshire, Pierre Ménard +1
We consider a stochastic bandit problem with a possibly infinite number of arms. We write for the proportion of optimal arms and for the minimal mean-gap between optimal…
Episodic Reinforcement Learning in Finite MDPs: Minimax Lower Bounds Revisited
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann +1
In this paper, we propose new problem-independent lower bounds on the sample complexity and regret in episodic MDPs, with a particular focus on the non-stationary case in which the…
Optimal Strategies for Graph-Structured Bandits
Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard
We study a structured variant of the multi-armed bandit problem specified by a set of Bernoulli distributions with mean…