3 papers
math.OC2025
Mirror Descent for Stochastic Control Problems with Measure-valued Controls
Bekzhan Kerimkulov, David Å iÅ¡ka, Åukasz Szpruch +1
This paper studies the convergence of the mirror descent algorithm for finite horizon stochastic control problems with measure-valued control processes. The control objective invol…
math.OC2025
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
Bekzhan Kerimkulov, James-Michael Leahy, David Siska +2
We study the global convergence of a Fisher-Rao policy gradient flow for infinite-horizon entropy-regularised Markov decision processes with Polish state and action space. The flow…
math.OC2025
Entropy annealing for policy mirror descent in continuous time and space
Deven Sethi, David Šiška, Yufei Zhang
Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additi…