23 citations · 38 across the 14 of their papers we have counts for
12 papers · 1 filter
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang +12
Mirror descent value iteration (MDVI), an abstraction of Kullback-Leibler (KL) and entropy-regularized reinforcement learning (RL), has served as the basis for recent high-performi…
Learning Generative Models with Goal-conditioned Reinforcement Learning
Mariana Vargas Vieyra, Pierre Ménard
We present a novel, alternative framework for learning generative models with goal-conditioned reinforcement learning. We define two agents, a goal conditioned agent (GC-agent) and…
KL-Entropy-Regularized RL with a Generative Model is Minimax Optimal
Tadashi Kozuno, Wenhao Yang, Nino Vieillard +10
In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model. Particularly, we analyze mirror descent value iteration (M…
Adaptive Multi-Goal Exploration
Jean Tarbouriech, Omar Darwiche Domingues, Pierre Ménard +3
We introduce a generic strategy for provably efficient multi-goal exploration. It relies on AdaGoal, a novel goal selection scheme that leverages a measure of uncertainty in reachi…
Bandits with many optimal arms
Rianne de Heide, James Cheshire, Pierre Ménard +1
We consider a stochastic bandit problem with a possibly infinite number of arms. We write for the proportion of optimal arms and for the minimal mean-gap between optimal…
Episodic Reinforcement Learning in Finite MDPs: Minimax Lower Bounds Revisited
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann +1
In this paper, we propose new problem-independent lower bounds on the sample complexity and regret in episodic MDPs, with a particular focus on the non-stationary case in which the…