1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 1 cited
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
Yuhao Ding, Junzi Zhang, Hyunin Lee +1
Entropy regularization is an efficient technique for encouraging exploration and preventing a premature convergence of (vanilla) policy gradient methods in reinforcement learning (…
cs.LG2021
On the Global Optimum Convergence of Momentum-based Policy Gradient
Yuhao Ding, Junzi Zhang, Javad Lavaei
Policy gradient (PG) methods are popular and efficient for large-scale reinforcement learning due to their relative stability and incremental nature. In recent years, the empirical…