8 papers
Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs
Ziyue Chen, David Šiška, Lukasz Szpruch
We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider lo…
A note on convergence of Wasserstein policy optimization
David Šiška, Yufei Zhang
Wasserstein Policy Optimization (WPO) is a recently proposed reinforcement learning algorithm that leverages Wasserstein gradient flows to optimize stochastic policies in continuou…
PPO in the Fisher-Rao geometry
Razvan-Andrei Lascu, David Å iÅ¡ka, Åukasz Szpruch
Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and conve…
Mirror Descent for Stochastic Control Problems with Measure-valued Controls
Bekzhan Kerimkulov, David Å iÅ¡ka, Åukasz Szpruch +1
This paper studies the convergence of the mirror descent algorithm for finite horizon stochastic control problems with measure-valued control processes. The control objective invol…
Logarithmic regret in the ergodic Avellaneda-Stoikov market making model
Jialun Cao, David Šiška, Lukasz Szpruch +1
We analyse the regret arising from learning the price sensitivity parameter of liquidity takers in the ergodic version of the Avellaneda-Stoikov market making model. We show t…
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
Bekzhan Kerimkulov, James-Michael Leahy, David Siska +2
We study the global convergence of a Fisher-Rao policy gradient flow for infinite-horizon entropy-regularised Markov decision processes with Polish state and action space. The flow…