1 paper
Yasin Abbasi-Yadkori, Peter L. Bartlett, Victor Gabillon +2
We study bandit best-arm identification with arbitrary and potentially adversarial rewards. A simple random uniform learner obtains the optimal rate of error in the adversarial sce…