Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs
Ziyue Chen, David Šiška, Lukasz Szpruch
We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider lo…
cs.LG2026
A note on convergence of Wasserstein policy optimization
David Šiška, Yufei Zhang
Wasserstein Policy Optimization (WPO) is a recently proposed reinforcement learning algorithm that leverages Wasserstein gradient flows to optimize stochastic policies in continuou…
cs.LG2026
PPO in the Fisher-Rao geometry
Razvan-Andrei Lascu, David Å iÅ¡ka, Åukasz Szpruch
Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and conve…