1 paper
Hao Qin, Kwang-Sung Jun, Chicheng Zhang
We study K-armed bandit problems where the reward distributions of the arms are all supported on the [0,1] interval. It has been a challenge to design regret-efficient randomiz…