68 citations · 93 across the 3 of their papers we have counts for
3 papers
Algorithms for CVaR Optimization in MDPs
Yinlam Chow, Mohammad Ghavamzadeh
In many sequential decision-making problems we may want to manage risk by minimizing some measure of variability in costs in addition to minimizing a standard criterion. Conditiona…
A Dantzig Selector Approach to Temporal Difference Learning
Matthieu Geist, Bruno Scherrer, Alessandro Lazaric +1
LSTD is a popular algorithm for value function approximation. Whenever the number of features is larger than the number of samples, it must be paired with some form of regularizati…
Approximate Modified Policy Iteration
Bruno Scherrer, Victor Gabillon, Mohammad Ghavamzadeh +1
Modified policy iteration (MPI) is a dynamic programming (DP) algorithm that contains the two celebrated policy and value iteration methods. Despite its generality, MPI has not bee…