1 paper · 1 filter
Yasin Abbasi-Yadkori, Peter L. Bartlett, Victor Gabillon +2
We study bandit best-arm identification with arbitrary and potentially adversarial rewards. A simple random uniform learner obtains the optimal rate of error in the adversarial sce…