68 citations · 96 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2014
Variance-Constrained Actor-Critic Algorithms for Discounted and Average Reward MDPs
Prashanth L. A., Mohammad Ghavamzadeh
In many sequential decision-making problems we may want to manage risk by minimizing some measure of variability in rewards in addition to maximizing a standard criterion. Variance…
cs.LG2012★ 24 cited
A Dantzig Selector Approach to Temporal Difference Learning
Matthieu Geist, Bruno Scherrer, Alessandro Lazaric +1
LSTD is a popular algorithm for value function approximation. Whenever the number of features is larger than the number of samples, it must be paired with some form of regularizati…