Showing 2025Show all
2 papers · 1 filter
cs.LG2025
Policy Newton methods for Distortion Riskmetrics
Soumen Pachal, Mizhaan Prajit Maniyar, Prashanth L. A
We consider the problem of risk-sensitive control in a reinforcement learning (RL) framework. In particular, we aim to find a risk-optimal policy by maximizing the distortion riskm…
stat.ML2025
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…