6 papers
Local and adaptive mirror descents in extensive-form games
Côme Fiegel, Pierre Ménard, Tadashi Kozuno +3
We study how to learn -optimal strategies in zero-sum imperfect information games (IIG) with trajectory feedback. In this setting, players update their policies sequentially bas…
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang +12
Mirror descent value iteration (MDVI), an abstraction of Kullback-Leibler (KL) and entropy-regularized reinforcement learning (RL), has served as the basis for recent high-performi…
Sharp Deviations Bounds for Dirichlet Weighted Sums with Application to analysis of Bayesian algorithms
Denis Belomestny, Pierre Menard, Alexey Naumov +2
In this work, we derive sharp non-asymptotic deviation bounds for weighted sums of Dirichlet random variables. These bounds are based on a novel integral representation of the dens…
Learning Generative Models with Goal-conditioned Reinforcement Learning
Mariana Vargas Vieyra, Pierre Ménard
We present a novel, alternative framework for learning generative models with goal-conditioned reinforcement learning. We define two agents, a goal conditioned agent (GC-agent) and…
Fast Rates for Maximum Entropy Exploration
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +7
We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maxim…
Indexed Minimum Empirical Divergence for Unimodal Bandits
Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard
We consider a multi-armed bandit problem specified by a set of one-dimensional family exponential distributions endowed with a unimodal structure. We introduce IMED-UB, a algorithm…