Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…
stat.ML2025
Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
This paper introduces a general framework for risk-sensitive bandits that integrates the notions of risk-sensitive objectives by adopting a rich class of distortion riskmetrics. Th…