4 papers
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning
Jialun Cao, Fernando Acero, David Šiška +1
Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigoro…
Convergence of an actor-critic gradient flow for entropy regularised MDPs in general spaces
Denis Zorba, David Šiška, Lukasz Szpruch
We prove the stability and global convergence of a coupled actor-critic gradient flow for infinite-horizon and entropy-regularised Markov decision processes (MDPs) in continuous st…
Mirror descent actor-critic methods for entropy regularised MDPs in general spaces: stability and convergence
Denis Zorba, David Šiška, Lukasz Szpruch
We provide theoretical guarantees for convergence of discrete-time policy mirror descent with inexact advantage functions updated using temporal difference (TD) learning for entrop…
Mirror descent for constrained stochastic control problems
Deven Sethi, David Šiška
Mirror descent is a well established tool for solving convex optimization problems with convex constraints. This article introduces continuous-time mirror descent dynamics for appr…