68 citations · 156 across the 8 of their papers we have counts for
1 paper · 1 filter
Yinlam Chow, Mohammad Ghavamzadeh
In this paper, we show how a simulated Markov decision process (MDP) built by the so-called \emph{baseline} policies, can be used to compute a different policy, namely the \emph{si…