1 paper
Ziyue Chen, David Šiška, Lukasz Szpruch
We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider lo…