48 citations · 175 across the 15 of their papers we have counts for
5 papers · 1 filter
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang +1
We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-ba…
Towards a Theoretical Foundation of Policy Optimization for Learning Control Policies
Bin Hu, Kaiqing Zhang, Na Li +3
Gradient-based methods have been widely used for system design and optimization in diverse application domains. Recently, there has been a renewed interest in studying theoretical…
Globally Convergent Policy Search over Dynamic Filters for Output Estimation
Jack Umenberger, Max Simchowitz, Juan C. Perdomo +2
We introduce the first direct policy search algorithm which provably converges to the globally optimal filter for the classical problem of predicting the outputs…
Policy Optimization for Linear Control with Robustness Guarantee: Implicit Regularization and Global Convergence
Kaiqing Zhang, Bin Hu, Tamer Başar
Policy optimization (PO) is a key ingredient for reinforcement learning (RL). For control design, certain constraints are usually enforced on the policies to optimize, accounting f…
Global Convergence of Policy Gradient Methods to (Almost) Locally Optimal Policies
Kaiqing Zhang, Alec Koppel, Hao Zhu +1
Policy gradient (PG) methods are a widely used reinforcement learning methodology in many applications such as video games, autonomous driving, and robotics. In spite of its empiri…