67 citations · 332 across the 55 of their papers we have counts for
11 papers · 2 filters
Differentially Private Exploration in Reinforcement Learning with Linear Representation
Paul Luyo, Evrard Garcelon, Alessandro Lazaric +1
This paper studies privacy-preserving exploration in Markov Decision Processes (MDPs) with linear representation. We first consider the setting of linear-mixture MDPs (Ayoub et al.…
Top Ranking for Multi-Armed Bandit with Noisy Evaluations
Evrard Garcelon, Vashist Avadhanula, Alessandro Lazaric +1
We consider a multi-armed bandit setting where, at the beginning of each round, the learner receives noisy independent, and possibly biased, \emph{evaluations} of the true reward o…
Adaptive Multi-Goal Exploration
Jean Tarbouriech, Omar Darwiche Domingues, Pierre Ménard +3
We introduce a generic strategy for provably efficient multi-goal exploration. It relies on AdaGoal, a novel goal selection scheme that leverages a measure of uncertainty in reachi…
Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection
Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano +3
We study the role of the representation of state-action value functions in regret minimization in finite-horizon Markov Decision Processes (MDPs) with linear structure. We first de…
Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching
Pierre-Alexandre Kamienny, Jean Tarbouriech, Sylvain Lamprier +2
Learning meaningful behaviors in the absence of reward is a difficult problem in reinforcement learning. A desirable and challenging unsupervised objective is to learn a set of div…
A general sample complexity analysis of vanilla policy gradient
Rui Yuan, Robert M. Gower, Alessandro Lazaric
We adapt recent tools developed for the analysis of Stochastic Gradient Descent (SGD) in non-convex optimization to obtain convergence and sample complexity guarantees for the vani…