activity
20122024
most citedRisk-Aversion in Multi-armed Bandits

93 citations · 218 across the 12 of their papers we have counts for

collaborators

9 papers

cs.LG202221 cited

Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Denis Yarats, David Brandfonbrener, Hao Liu +4

Recent progress in deep learning has relied on access to large and diverse datasets. Such data-driven progress has been less evident in offline reinforcement learning (RL), because…

cs.LG20211 cited

Differentially Private Exploration in Reinforcement Learning with Linear Representation

Paul Luyo, Evrard Garcelon, Alessandro Lazaric +1

This paper studies privacy-preserving exploration in Markov Decision Processes (MDPs) with linear representation. We first consider the setting of linear-mixture MDPs (Ayoub et al.…

cs.LG2021

Top Ranking for Multi-Armed Bandit with Noisy Evaluations

Evrard Garcelon, Vashist Avadhanula, Alessandro Lazaric +1

We consider a multi-armed bandit setting where, at the beginning of each round, the learner receives noisy independent, and possibly biased, \emph{evaluations} of the true reward o…

stat.ML2016

Analysis of Kelner and Levin graph sparsification algorithm for a streaming setting

Daniele Calandriello, Alessandro Lazaric, Michal Valko

We derive a new proof to show that the incremental resparsification algorithm proposed by Kelner and Levin (2013) produces a spectral sparsifier in high probability. We rigorously…

cs.AI20163 cited

Open Problem: Approximate Planning of POMDPs in the class of Memoryless Policies

Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

Planning plays an important role in the broad class of decision theory. Planning has drawn much attention in recent work in the robotics and sequential decision making areas. Recen…

cs.LG201475 cited

Best-Arm Identification in Linear Bandits

Marta Soare, Alessandro Lazaric, Rémi Munos

We study the best-arm identification problem in linear bandit, where the rewards of the arms depend linearly on an unknown parameter and the objective is to return the arm wi…