activity
20162022
most citedFinite-Time Analysis of Asynchronous Stochastic Approximation and -Learning

24 citations · 77 across the 9 of their papers we have counts for

collaborators

15 papers

cs.LG20221 cited

Global Convergence of Localized Policy Iteration in Networked Multi-Agent Reinforcement Learning

Yizhou Zhang, Guannan Qu, Pan Xu +3

We study a multi-agent reinforcement learning (MARL) problem where the agents interact over a given network. The goal of the agents is to cooperatively maximize the average of thei…

math.OC20221 cited

Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and Nonlinearity

Yiheng Lin, Yang Hu, Guannan Qu +2

We study Model Predictive Control (MPC) and propose a general analysis pipeline to bound its dynamic regret. The pipeline first requires deriving a perturbation bound for a finite-…

math.OC20224 cited

On the Sample Complexity of Stabilizing LTI Systems on a Single Trajectory

Yang Hu, Adam Wierman, Guannan Qu

Stabilizing an unknown dynamical system is one of the central problems in control theory. In this paper, we study the sample complexity of the learn-to-stabilize problem in Linear…

cs.MA2021

Decentralized Graph-Based Multi-Agent Reinforcement Learning Using Reward Machines

Jueming Hu, Zhe Xu, Weichang Wang +3

In multi-agent reinforcement learning (MARL), it is challenging for a collection of agents to learn complex temporally extended tasks. The difficulties lie in computational complex…

eess.SY20213 cited

Stability Constrained Reinforcement Learning for Real-Time Voltage Control

Yuanyuan Shi, Guannan Qu, Steven Low +2

Deep reinforcement learning (RL) has been recognized as a promising tool to address the challenges in real-time control of power systems. However, its deployment in real-world powe…

math.OC202110 cited

Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying Systems

Yiheng Lin, Yang Hu, Haoyuan Sun +3

We study predictive control in a setting where the dynamics are time-varying and linear, and the costs are time-varying and well-conditioned. At each time step, the controller rece…