6 papers
Gradient-based Stochastic Optimization of Utility-based Shortfall Risk
Sumedh Gupte, Prashanth L. A., Sanjay P. Bhat
We consider the problems of estimation and optimization of utility-based shortfall risk (UBSR). We extend UBSR to cover possibly unbounded random variables. We cover prominent risk…
Generalized Random Direction Newton Algorithms for Stochastic Optimization
Soumen Pachal, Prashanth L. A., Shalabh Bhatnagar +1
We present a family of generalized Hessian estimators of the objective using random direction stochastic approximation (RDSA) by utilizing only noisy function measurements. The for…
Policy Newton methods for Distortion Riskmetrics
Soumen Pachal, Mizhaan Prajit Maniyar, Prashanth L. A
We consider the problem of risk-sensitive control in a reinforcement learning (RL) framework. In particular, we aim to find a risk-optimal policy by maximizing the distortion riskm…
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…
A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP
Tejaram Sangadi, L. A. Prashanth, Krishna Jagannathan
Motivated by applications in risk-sensitive reinforcement learning, we study mean-variance optimization in a discounted reward Markov Decision Process (MDP). Specifically, we analy…
Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms
Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2
This paper introduces a general framework for risk-sensitive bandits that integrates the notions of risk-sensitive objectives by adopting a rich class of distortion riskmetrics. Th…