107 citations
- Centre National de la Recherche ScientifiqueFR6 papers
- Google DeepMind (United Kingdom)GB6 papers
- Centre de Recherche en Informatique, Signal et Automatique de LilleFR4 papers
- Meta (Israel)IL4 papers
- Google (United States)US3 papers
- Centre de Recherche en InformatiqueFR2 papers
- Département d'InformatiqueFR2 papers
- École Centrale de LilleFR2 papers
- École Normale Supérieure - PSLFR2 papers
- Huawei Technologies (United Kingdom)GB2 papers
- Sequella (United States)US2 papers
- Université de LilleFR2 papers
7 papers · 1 filter
Improved Sample Complexity for Incremental Autonomous Exploration in MDPs
Jean Tarbouriech, Matteo Pirotta, Michal Valko +1
We investigate the exploration of an unknown environment when no reward function is provided. Building on the incremental exploration setting introduced by Lim and Auer [1], we def…
Self-Imitation Advantage Learning
Johan Ferret, Olivier Pietquin, Matthieu Geist
Self-imitation learning is a Reinforcement Learning (RL) method that encourages actions whose returns were higher than expected, which helps in hard exploration and sparse reward p…
Optimal Strategies for Graph-Structured Bandits
Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard
We study a structured variant of the multi-armed bandit problem specified by a set of Bernoulli distributions with mean…
Gamification of Pure Exploration for Linear Bandits
Rémy Degenne, Pierre Ménard, Xuedong Shang +1
We investigate an active pure-exploration setting, that includes best-arm identification, in the context of linear stochastic bandits. While asymptotically optimal algorithms exist…
Forced-exploration free Strategies for Unimodal Bandits
Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard
We consider a multi-armed bandit problem specified by a set of Gaussian or Bernoulli distributions endowed with a unimodal structure. Although this problem has been addressed in th…
What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk +9
In recent years, on-policy reinforcement learning (RL) has been successfully applied to many different continuous control tasks. While RL algorithms are often conceptually simple,…