activity
20242026
collaborators

8 papers

cs.LG2026

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

Ziyue Chen, David Šiška, Lukasz Szpruch

We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider lo…

cs.LG2026

A note on convergence of Wasserstein policy optimization

David Šiška, Yufei Zhang

Wasserstein Policy Optimization (WPO) is a recently proposed reinforcement learning algorithm that leverages Wasserstein gradient flows to optimize stochastic policies in continuou…

cs.LG2026

PPO in the Fisher-Rao geometry

Razvan-Andrei Lascu, David Šiška, Łukasz Szpruch

Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and conve…

math.OC2025

Mirror Descent for Stochastic Control Problems with Measure-valued Controls

Bekzhan Kerimkulov, David Šiška, Łukasz Szpruch +1

This paper studies the convergence of the mirror descent algorithm for finite horizon stochastic control problems with measure-valued control processes. The control objective invol…

math.OC2025

Logarithmic regret in the ergodic Avellaneda-Stoikov market making model

Jialun Cao, David Šiška, Lukasz Szpruch +1

We analyse the regret arising from learning the price sensitivity parameter of liquidity takers in the ergodic version of the Avellaneda-Stoikov market making model. We show t…

math.OC2025

A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces

Bekzhan Kerimkulov, James-Michael Leahy, David Siska +2

We study the global convergence of a Fisher-Rao policy gradient flow for infinite-horizon entropy-regularised Markov decision processes with Polish state and action space. The flow…