3 papers
cs.LG2024
Dynamic Learning Rate for Deep Reinforcement Learning: A Bandit Approach
Henrique Donâncio, Antoine Barrier, Leah F. South +1
In deep Reinforcement Learning (RL), the learning rate critically influences both stability and performance, yet its optimal value shifts during training as the environment and pol…
cs.LG2022
On Best-Arm Identification with a Fixed Budget in Non-Parametric Multi-Armed Bandits
Antoine Barrier, Aurélien Garivier, Gilles Stoltz
We lay the foundations of a non-parametric theory of best-arm identification in multi-armed bandits with a fixed budget T. We consider general, possibly non-parametric, models D fo…
math.ST2021
A Non-asymptotic Approach to Best-Arm Identification for Gaussian Bandits
Antoine Barrier, Aurélien Garivier, Tomáš Kocák
We propose a new strategy for best-arm identification with fixed confidence of Gaussian variables with bounded means and unit variance. This strategy, called Exploration-Biased Sam…