29 citations · 67 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2020★ 14 cited
Single-Timescale Stochastic Nonconvex-Concave Optimization for Smooth Nonlinear TD Learning
Shuang Qiu, Zhuoran Yang, Xiaohan Wei +2
Temporal-Difference (TD) learning with nonlinear smooth function approximation for policy evaluation has achieved great success in modern reinforcement learning. It is shown that s…
cs.LG2020
Provably Efficient Safe Exploration via Primal-Dual Policy Optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang +2
We study the Safe Reinforcement Learning (SRL) problem using the Constrained Markov Decision Process (CMDP) formulation in which an agent aims to maximize the expected total reward…