output
20192026
most citedWhat Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

107 citations

Showing 2020Show all

7 papers · 1 filter

cs.LG20209 cited

Improved Sample Complexity for Incremental Autonomous Exploration in MDPs

Jean Tarbouriech, Matteo Pirotta, Michal Valko +1

We investigate the exploration of an unknown environment when no reward function is provided. Building on the incremental exploration setting introduced by Lim and Auer [1], we def…

cs.LG20208 cited

Self-Imitation Advantage Learning

Johan Ferret, Olivier Pietquin, Matthieu Geist

Self-imitation learning is a Reinforcement Learning (RL) method that encourages actions whose returns were higher than expected, which helps in hard exploration and sparse reward p…

cs.IT2020

Optimal Strategies for Graph-Structured Bandits

Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard

We study a structured variant of the multi-armed bandit problem specified by a set of Bernoulli distributions with mean…

stat.ML202023 cited

Gamification of Pure Exploration for Linear Bandits

Rémy Degenne, Pierre Ménard, Xuedong Shang +1

We investigate an active pure-exploration setting, that includes best-arm identification, in the context of linear stochastic bandits. While asymptotically optimal algorithms exist…

cs.LG20203 cited

Forced-exploration free Strategies for Unimodal Bandits

Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard

We consider a multi-armed bandit problem specified by a set of Gaussian or Bernoulli distributions endowed with a unimodal structure. Although this problem has been addressed in th…

cs.LG2020107 cited

What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk +9

In recent years, on-policy reinforcement learning (RL) has been successfully applied to many different continuous control tasks. While RL algorithms are often conceptually simple,…