activity
20192022
most citedScalable Multi-Agent Reinforcement Learning for Networked Systems with Average Reward

23 citations · 35 across the 6 of their papers we have counts for

collaborators

8 papers

cs.LG20221 cited

Global Convergence of Localized Policy Iteration in Networked Multi-Agent Reinforcement Learning

Yizhou Zhang, Guannan Qu, Pan Xu +3

We study a multi-agent reinforcement learning (MARL) problem where the agents interact over a given network. The goal of the agents is to cooperatively maximize the average of thei…

math.OC20221 cited

Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and Nonlinearity

Yiheng Lin, Yang Hu, Guannan Qu +2

We study Model Predictive Control (MPC) and propose a general analysis pipeline to bound its dynamic regret. The pipeline first requires deriving a perturbation bound for a finite-…

cs.LG2021

Online Optimization with Feedback Delay and Nonlinear Switching Cost

Weici Pan, Guanya Shi, Yiheng Lin +1

We study a variant of online optimization in which the learner receives -round about hitting cost and there is a multi-step nonlinear switching cost,…

math.OC202110 cited

Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying Systems

Yiheng Lin, Yang Hu, Haoyuan Sun +3

We study predictive control in a setting where the dynamics are time-varying and linear, and the costs are time-varying and well-conditioned. At each time step, the controller rece…

math.OC202023 cited

Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average Reward

Guannan Qu, Yiheng Lin, Adam Wierman +1

It has long been recognized that multi-agent reinforcement learning (MARL) faces significant scalability issues due to the fact that the size of the state and action spaces are exp…

cs.LG2020

Online Optimization with Memory and Competitive Control

Guanya Shi, Yiheng Lin, Soon-Jo Chung +2

This paper presents competitive algorithms for a novel class of online optimization problems with memory. We consider a setting where the learner seeks to minimize the sum of a hit…