29 citations · 67 across the 6 of their papers we have counts for
Showing 2020Show all
3 papers · 1 filter
cs.LG2020★ 14 cited
Single-Timescale Stochastic Nonconvex-Concave Optimization for Smooth Nonlinear TD Learning
Shuang Qiu, Zhuoran Yang, Xiaohan Wei +2
Temporal-Difference (TD) learning with nonlinear smooth function approximation for policy evaluation has achieved great success in modern reinforcement learning. It is shown that s…
math.ST2020
II. High Dimensional Estimation under Weak Moment Assumptions: Structured Recovery and Matrix Estimation
Xiaohan Wei
The purpose of this thesis is to develop new theories on high-dimensional structured signal recovery under a rather weak assumption on the measurements that only a finite number of…
cs.LG2020
Provably Efficient Safe Exploration via Primal-Dual Policy Optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang +2
We study the Safe Reinforcement Learning (SRL) problem using the Constrained Markov Decision Process (CMDP) formulation in which an agent aims to maximize the expected total reward…