4 papers · 1 filter
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…
Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
This paper introduces a general framework for risk-sensitive bandits that integrates the notions of risk-sensitive objectives by adopting a rich class of distortion riskmetrics. Th…
Optimal Best Arm Identification with Fixed Confidence in Restless Bandits
P. N. Karthik, Vincent Y. F. Tan, Arpan Mukherjee +1
We study best arm identification in a restless multi-armed bandit setting with finitely many arms. The discrete-time data generated by each arm forms a homogeneous Markov chain tak…
Improved Bound for Robust Causal Bandits with Linear Models
Zirui Yan, Arpan Mukherjee, Burak Varıcı +1
This paper investigates the robustness of causal bandits (CBs) in the face of temporal model fluctuations. This setting deviates from the existing literature's widely-adopted assum…