1 paper · 1 filter
Subhodip Panda, Shubhada Agrawal
We study the tail behavior of regret in stochastic multi-armed bandits for algorithms that are asymptotically optimal in expectation. While minimizing expected regret is the classi…