activity
20152021
most citedOn the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost

27 citations · 150 across the 13 of their papers we have counts for

collaborators

29 papers

cs.LG2021

Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality

Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1

Designing off-policy reinforcement learning algorithms is typically a very challenging task, because a desirable iteration update often involves an expectation over an on-policy di…

cs.LG2020

Provable Fictitious Play for General Mean-Field Games

Qiaomin Xie, Zhuoran Yang, Zhaoran Wang +1

We propose a reinforcement learning algorithm for stationary mean-field games, where the goal is to learn a pair of mean-field state and stationary policy that constitutes the Nash…

math.OC20206 cited

Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous Time

Weichen Wang, Jiequn Han, Zhuoran Yang +1

Reinforcement learning is a powerful tool to learn the optimal policy of possibly multiple agents by interacting with the environment. As the number of agents grow to be very large…

stat.ML202016 cited

Accelerating Nonconvex Learning via Replica Exchange Langevin Diffusion

Yi Chen, Jinglin Chen, Jing Dong +2

Langevin diffusion is a powerful method for nonconvex optimization, which enables the escape from local minima by injecting noise into the gradient. In particular, the temperature…

stat.ML2020

Provably Efficient Neural Estimation of Structural Equation Model: An Adversarial Approach

Luofeng Liao, You-Lin Chen, Zhuoran Yang +3

Structural equation models (SEMs) are widely used in sciences, ranging from economics to psychology, to uncover causal relationships underlying a complex system under consideration…

cs.LG20207 cited

On the Global Optimality of Model-Agnostic Meta-Learning

Lingxiao Wang, Qi Cai, Zhuoran Yang +1

Model-agnostic meta-learning (MAML) formulates meta-learning as a bilevel optimization problem, where the inner level solves each subtask based on a shared prior, while the outer l…