activity
20092022
most citedAn Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient Methods

30 citations · 106 across the 23 of their papers we have counts for

collaborators
Showing 2019Show all

16 papers · 1 filter

cs.LG201919 cited

Decentralized Multi-Agent Reinforcement Learning with Networked Agents: Recent Advances

Kaiqing Zhang, Zhuoran Yang, Tamer Başar

Multi-agent reinforcement learning (MARL) has long been a significant and everlasting research topic in both machine learning and control. With the recent development of (single-ag…

math.OC2019

Zero-Sum Differential Games on the Wasserstein Space

Jun Moon, Tamer Basar

We consider two-player zero-sum differential games (ZSDGs), where the state process (dynamical system) depends on the random initial condition and the state process's distribution,…

cs.GT201915 cited

Non-Cooperative Inverse Reinforcement Learning

Xiangyuan Zhang, Kaiqing Zhang, Erik Miehling +1

Making decisions in the presence of a strategic opponent requires one to take into account the opponent's ability to actively mask its intended objective. To describe such strategi…

math.OC2019

Quantifying Market Efficiency Impacts of Aggregated Distributed Energy Resources

Khaled Alshehri, Mariola Ndrio, Subhonmesh Bose +1

We focus on the aggregation of distributed energy resources (DERs) through a profit-maximizing intermediary that enables participation of DERs in wholesale electricity markets. Par…

math.OC2019

Policy Optimization for Linear Control with Robustness Guarantee: Implicit Regularization and Global Convergence

Kaiqing Zhang, Bin Hu, Tamer Başar

Policy optimization (PO) is a key ingredient for reinforcement learning (RL). For control design, certain constraints are usually enforced on the policies to optimize, accounting f…

cs.GT2019

Strategic Inference with a Single Private Sample

Erik Miehling, Roy Dong, Cédric Langbort +1

Motivated by applications in cyber security, we develop a simple game model for describing how a learning agent's private information influences an observing agent's inference proc…