Showing stat.MLShow all
2 papers · 1 filter
stat.ML2024
Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more
Nobuaki Kikkawa, Hiroshi Ohno
The upper confidence bound (UCB) policy is recognized as an order-optimal solution for the classical total-reward bandit problem. While similar UCB-based approaches have been appli…
stat.ML2022
Materials Discovery using Max K-Armed Bandit
Nobuaki Kikkawa, Hiroshi Ohno
Search algorithms for the bandit problems are applicable in materials discovery. However, the objectives of the conventional bandit problem are different from those of materials di…