2 papers
cs.LG2025
A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP
Tejaram Sangadi, L. A. Prashanth, Krishna Jagannathan
Motivated by applications in risk-sensitive reinforcement learning, we study mean-variance optimization in a discounted reward Markov Decision Process (MDP). Specifically, we analy…
cs.LG2025
Fixed-Confidence Best Arm Identification with Decreasing Variance
Tamojeet Roychowdhury, Kota Srinivas Reddy, Krishna P Jagannathan +1
We focus on the problem of best-arm identification in a stochastic multi-arm bandit with temporally decreasing variances for the arms' rewards. We model arm rewards as Gaussian ran…