activity
20152023
most citedLearning Safe Multi-Agent Control with Decentralized Neural Barrier Certificates

48 citations · 175 across the 15 of their papers we have counts for

collaborators
Showing math.OCShow all

5 papers · 1 filter

math.OC2023

Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs

Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang +1

We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-ba…

math.OC20226 cited

Towards a Theoretical Foundation of Policy Optimization for Learning Control Policies

Bin Hu, Kaiqing Zhang, Na Li +3

Gradient-based methods have been widely used for system design and optimization in diverse application domains. Recently, there has been a renewed interest in studying theoretical…

math.OC20225 cited

Globally Convergent Policy Search over Dynamic Filters for Output Estimation

Jack Umenberger, Max Simchowitz, Juan C. Perdomo +2

We introduce the first direct policy search algorithm which provably converges to the globally optimal filter for the classical problem of predicting the outputs…

math.OC2019

Policy Optimization for Linear Control with Robustness Guarantee: Implicit Regularization and Global Convergence

Kaiqing Zhang, Bin Hu, Tamer Başar

Policy optimization (PO) is a key ingredient for reinforcement learning (RL). For control design, certain constraints are usually enforced on the policies to optimize, accounting f…

math.OC2019

Global Convergence of Policy Gradient Methods to (Almost) Locally Optimal Policies

Kaiqing Zhang, Alec Koppel, Hao Zhu +1

Policy gradient (PG) methods are a widely used reinforcement learning methodology in many applications such as video games, autonomous driving, and robotics. In spite of its empiri…