1 paper
Antoine Salomon, Jean-Yves Audibert
This paper studies the deviations of the regret in a stochastic multi-armed bandit problem. When the total number of plays n is known beforehand by the agent, Audibert et al. (2009…