4 papers
Policy Newton methods for Distortion Riskmetrics
Soumen Pachal, Mizhaan Prajit Maniyar, Prashanth L. A
We consider the problem of risk-sensitive control in a reinforcement learning (RL) framework. In particular, we aim to find a risk-optimal policy by maximizing the distortion riskm…
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…
A Survey of Risk-Aware Multi-Armed Bandits
Vincent Y. F. Tan, Prashanth L. A., Krishna Jagannathan
In several applications such as clinical trials and financial portfolio optimization, the expected value (or the average reward) does not satisfactorily capture the merits of a dru…
Estimation of Spectral Risk Measures
Ajay Kumar Pandey, Prashanth L. A., Sanjay P. Bhat
We consider the problem of estimating a spectral risk measure (SRM) from i.i.d. samples, and propose a novel method that is based on numerical integration. We show that our SRM est…