64 citations · 208 across the 15 of their papers we have counts for
5 papers · 1 filter
Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement Learning
Stefan Stojanovic, Yassir Jedra, Alexandre Proutiere
We study matrix estimation problems arising in reinforcement learning (RL) with low-rank structure. In low-rank bandits, the matrix to be recovered specifies the expected arm rewar…
Best Policy Identification in Linear MDPs
Jerome Taupin, Yassir Jedra, Alexandre Proutiere
We investigate the problem of best policy identification in discounted linear Markov Decision Processes in the fixed confidence setting under a generative model. We first derive an…
Unimodal Bandits without Smoothness
Richard Combes, Alexandre Proutiere
We consider stochastic bandit problems with a continuous set of arms and where the expected reward is a continuous and unimodal function of the arm. No further assumption is made r…
Unimodal Bandits: Regret Lower Bounds and Optimal Algorithms
Richard Combes, Alexandre Proutiere
We consider stochastic multi-armed bandits where the expected reward is a unimodal function over partially ordered arms. This important class of problems has been recently investig…
Lipschitz Bandits: Regret Lower Bounds and Optimal Algorithms
Stefan Magureanu, Richard Combes, Alexandre Proutiere
We consider stochastic multi-armed bandit problems where the expected reward is a Lipschitz function of the arm, and where the set of arms is either discrete or continuous. For dis…