Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…
stat.ML2022
A Survey of Risk-Aware Multi-Armed Bandits
Vincent Y. F. Tan, Prashanth L. A., Krishna Jagannathan
In several applications such as clinical trials and financial portfolio optimization, the expected value (or the average reward) does not satisfactorily capture the merits of a dru…