activity
20152026
most citedMastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

67 citations · 332 across the 55 of their papers we have counts for

collaborators
Showing 2021 · cs.LGShow all

11 papers · 2 filters

cs.LG2021★ 1 cited

Differentially Private Exploration in Reinforcement Learning with Linear Representation

Paul Luyo, Evrard Garcelon, Alessandro Lazaric +1

This paper studies privacy-preserving exploration in Markov Decision Processes (MDPs) with linear representation. We first consider the setting of linear-mixture MDPs (Ayoub et al.…

cs.LG2021

Top Ranking for Multi-Armed Bandit with Noisy Evaluations

Evrard Garcelon, Vashist Avadhanula, Alessandro Lazaric +1

We consider a multi-armed bandit setting where, at the beginning of each round, the learner receives noisy independent, and possibly biased, \emph{evaluations} of the true reward o…

cs.LG2021

Adaptive Multi-Goal Exploration

Jean Tarbouriech, Omar Darwiche Domingues, Pierre Ménard +3

We introduce a generic strategy for provably efficient multi-goal exploration. It relies on AdaGoal, a novel goal selection scheme that leverages a measure of uncertainty in reachi…

cs.LG2021★ 2 cited

Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection

Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano +3

We study the role of the representation of state-action value functions in regret minimization in finite-horizon Markov Decision Processes (MDPs) with linear structure. We first de…

cs.LG2021

Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching

Pierre-Alexandre Kamienny, Jean Tarbouriech, Sylvain Lamprier +2

Learning meaningful behaviors in the absence of reward is a difficult problem in reinforcement learning. A desirable and challenging unsupervised objective is to learn a set of div…

cs.LG2021★ 8 cited

A general sample complexity analysis of vanilla policy gradient

Rui Yuan, Robert M. Gower, Alessandro Lazaric

We adapt recent tools developed for the analysis of Stochastic Gradient Descent (SGD) in non-convex optimization to obtain convergence and sample complexity guarantees for the vani…