3 papers
cs.LG2020
On Convergence of Gradient Expected Sarsa()
Long Yang, Gang Zheng, Yu Zhang +3
We study the convergence of with linear function approximation. We show that applying the off-line estimate (multi-step bootstrapping) to $\mathtt{Expe…
cs.LG2019
FiDi-RL: Incorporating Deep Reinforcement Learning with Finite-Difference Policy Search for Efficient Learning of Continuous Control
Longxiang Shi, Shijian Li, Longbing Cao +3
In recent years significant progress has been made in dealing with challenging problems using reinforcement learning.Despite its great success, reinforcement learning still faces c…
cs.LG2019
Policy Optimization with Stochastic Mirror Descent
Long Yang, Yu Zhang, Gang Zheng +5
Improving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes algorithm: a sample efficient policy gradient method with s…