1 paper
Akram Erraqabi, Alessandro Lazaric, Michal Valko +2
In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm…