42 citations · 42 across the 1 of their papers we have counts for
1 paper
Yichen Chen, Mengdi Wang
We study the online estimation of the optimal policy of a Markov decision process (MDP). We propose a class of Stochastic Primal-Dual (SPD) methods which exploit the inherent minim…