43 citations · 247 across the 30 of their papers we have counts for
5 papers · 1 filter
Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
Wanxin Jin, Zhaoran Wang, Zhuoran Yang +1
This paper develops a Pontryagin Differentiable Programming (PDP) methodology, which establishes a unified framework to solve a broad class of learning and control tasks. The PDP d…
Convergent Policy Optimization for Safe Reinforcement Learning
Ming Yu, Zhuoran Yang, Mladen Kolar +1
We study the safe reinforcement learning problem with nonlinear function approximation, where policy optimization is formulated as a constrained optimization problem with both the…
Provably Efficient Reinforcement Learning with Linear Function Approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang +1
Modern Reinforcement Learning (RL) is commonly applied to practical problems with an enormous number of states, where function approximation must be deployed to approximate either…
Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima
Qi Cai, Zhuoran Yang, Jason D. Lee +1
Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in v…
A Multi-Agent Off-Policy Actor-Critic Algorithm for Distributed Reinforcement Learning
Wesley Suttle, Zhuoran Yang, Kaiqing Zhang +3
This paper extends off-policy reinforcement learning to the multi-agent case in which a set of networked agents communicating with their neighbors according to a time-varying graph…