2 papers
cs.LG2026
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
Michael Lu, Max Qiushi Lin, Mo Chen +1
We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture po…
cs.LG2025
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
Max Qiushi Lin, Jincheng Mei, Matin Aghaei +6
Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typica…