3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.LG2019★ 1 cited
Correlated bandits or: How to minimize mean-squared error online
Vinay Praneeth Boda, Prashanth L. A
While the objective in traditional multi-armed bandit problems is to find the arm with the highest mean, in many settings, finding an arm that best captures information about other…
math.OC2014★ 3 cited
Simultaneous Perturbation Algorithms for Batch Off-Policy Search
Raphael Fonteneau, L. A. Prashanth
We propose novel policy search algorithms in the context of off-policy, batch mode reinforcement learning (RL) with continuous state and action spaces. Given a batch collection of…
cs.LG2014
Variance-Constrained Actor-Critic Algorithms for Discounted and Average Reward MDPs
Prashanth L. A., Mohammad Ghavamzadeh
In many sequential decision-making problems we may want to manage risk by minimizing some measure of variability in rewards in addition to maximizing a standard criterion. Variance…