2 papers
math.OC2026
Convergence of an actor-critic gradient flow for entropy regularised MDPs in general spaces
Denis Zorba, David Šiška, Lukasz Szpruch
We prove the stability and global convergence of a coupled actor-critic gradient flow for infinite-horizon and entropy-regularised Markov decision processes (MDPs) in continuous st…
math.OC2026
Mirror descent actor-critic methods for entropy regularised MDPs in general spaces: stability and convergence
Denis Zorba, David Šiška, Lukasz Szpruch
We provide theoretical guarantees for convergence of discrete-time policy mirror descent with inexact advantage functions updated using temporal difference (TD) learning for entrop…