collaborators

6 papers

cs.CE2026

Gradient-based Stochastic Optimization of Utility-based Shortfall Risk

Sumedh Gupte, Prashanth L. A., Sanjay P. Bhat

We consider the problems of estimation and optimization of utility-based shortfall risk (UBSR). We extend UBSR to cover possibly unbounded random variables. We cover prominent risk…

cs.LG2026

Generalized Random Direction Newton Algorithms for Stochastic Optimization

Soumen Pachal, Prashanth L. A., Shalabh Bhatnagar +1

We present a family of generalized Hessian estimators of the objective using random direction stochastic approximation (RDSA) by utilizing only noisy function measurements. The for…

cs.LG2025

Policy Newton methods for Distortion Riskmetrics

Soumen Pachal, Mizhaan Prajit Maniyar, Prashanth L. A

We consider the problem of risk-sensitive control in a reinforcement learning (RL) framework. In particular, we aim to find a risk-optimal policy by maximizing the distortion riskm…

stat.ML2025

Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms

Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2

The objective of canonical multi-armed bandits is to identify and repeatedly select an arm with the largest reward, often in the form of the expected value of the arm's probability…

cs.LG2025

A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP

Tejaram Sangadi, L. A. Prashanth, Krishna Jagannathan

Motivated by applications in risk-sensitive reinforcement learning, we study mean-variance optimization in a discounted reward Markov Decision Process (MDP). Specifically, we analy…

stat.ML2025

Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms

Meltem Tatlı, Arpan Mukherjee, Prashanth L. A. +2

This paper introduces a general framework for risk-sensitive bandits that integrates the notions of risk-sensitive objectives by adopting a rich class of distortion riskmetrics. Th…