1 paper · 1 filter
Nobuaki Kikkawa, Hiroshi Ohno
The upper confidence bound (UCB) policy is recognized as an order-optimal solution for the classical total-reward bandit problem. While similar UCB-based approaches have been appli…