3 citations · 3 across the 1 of their papers we have counts for
3 papers
cs.LG2019★ 3 cited
Multi-Path Policy Optimization
Ling Pan, Qingpeng Cai, Longbo Huang
Recent years have witnessed a tremendous improvement of deep reinforcement learning. However, a challenging problem is that an agent may suffer from inefficient exploration, partic…
cs.LG2019
Deterministic Value-Policy Gradients
Qingpeng Cai, Ling Pan, Pingzhong Tang
Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG alg…
cs.LG2019
Reinforcement Learning with Dynamic Boltzmann Softmax Updates
Ling Pan, Qingpeng Cai, Qi Meng +3
Value function estimation is an important task in reinforcement learning, i.e., prediction. The Boltzmann softmax operator is a natural value estimator and can provide several bene…