paper

Stochastic Multi-armed Bandits in Constant Space

arXiv:1712.09007

Abstract

We consider the stochastic bandit problem in the sublinear space setting, where one cannot record the win-loss record for all arms. We give an algorithm using words of space with regret \[ \sum_{i=1}^{K}\frac{1}{Δ_i}\log \frac{Δ_i}Δ\log T \] where is the gap between the best arm and arm and is the gap between the best and the second-best arms. If the rewards are bounded away from and , this is within an factor of the optimum regret possible without space constraints.

AISTATS 2018