activity
20172025
most citedNeural Certificates for Safe Control Policies

43 citations · 392 across the 81 of their papers we have counts for

collaborators
Showing 2019Show all

21 papers · 1 filter

cs.LG2019★ 19 cited

Decentralized Multi-Agent Reinforcement Learning with Networked Agents: Recent Advances

Kaiqing Zhang, Zhuoran Yang, Tamer Başar

Multi-agent reinforcement learning (MARL) has long been a significant and everlasting research topic in both machine learning and control. With the recent development of (single-ag…

cs.LG2019

Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework

Wanxin Jin, Zhaoran Wang, Zhuoran Yang +1

This paper develops a Pontryagin Differentiable Programming (PDP) methodology, which establishes a unified framework to solve a broad class of learning and control tasks. The PDP d…

cs.LG2019

Natural Actor-Critic Converges Globally for Hierarchical Linear Quadratic Regulator

Yuwei Luo, Zhuoran Yang, Zhaoran Wang +1

Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algo…

cs.LG2019

Provably Efficient Exploration in Policy Optimization

Qi Cai, Zhuoran Yang, Chi Jin +1

While policy-based reinforcement learning (RL) achieves tremendous successes in practice, it is significantly less understood in theory, especially compared with value-based RL. In…

cs.LG2019

Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms

Kaiqing Zhang, Zhuoran Yang, Tamer Başar

Recent years have witnessed significant advances in reinforcement learning (RL), which has registered great success in solving various sequential decision-making problems in machin…

cs.LG2019★ 30 cited

Convergent Policy Optimization for Safe Reinforcement Learning

Ming Yu, Zhuoran Yang, Mladen Kolar +1

We study the safe reinforcement learning problem with nonlinear function approximation, where policy optimization is formulated as a constrained optimization problem with both the…