1 citations · 1 across the 1 of their papers we have counts for
2 papers
math.OC2021
A Stochastic Composite Augmented Lagrangian Method For Reinforcement Learning
Yongfeng Li, Mingming Zhao, Weijie Chen +1
In this paper, we consider the linear programming (LP) formulation for deep reinforcement learning. The number of the constraints depends on the size of state and action spaces, wh…
math.OC2019★ 1 cited
A Stochastic Trust-Region Framework for Policy Optimization
Mingming Zhao, Yongfeng Li, Zaiwen Wen
In this paper, we study a few challenging theoretical and numerical issues on the well known trust region policy optimization for deep reinforcement learning. The goal is to find a…