1 citations · 1 across the 3 of their papers we have counts for
4 papers
Robust Action Gap Increasing with Clipped Advantage Learning
Zhe Zhang, Yaozhong Gan, Xiaoyang Tan
Advantage Learning (AL) seeks to increase the action gap between the optimal action and its competitors, so as to improve the robustness to estimation errors. However, the method b…
Smoothing Advantage Learning
Yaozhong Gan, Zhe Zhang, Xiaoyang Tan
Advantage learning (AL) aims to improve the robustness of value-based reinforcement learning against estimation errors with action-gap-based regularization. Unfortunately, the meth…
Stabilizing Q Learning Via Soft Mellowmax Operator
Yaozhong Gan, Zhe Zhang, Xiaoyang Tan
Learning complicated value functions in high dimensional state space by function approximation is a challenging task, partially due to that the max-operator used in temporal differ…
Trust Region-Guided Proximal Policy Optimization
Yuhui Wang, Hao He, Xiaoyang Tan +1
Proximal policy optimization (PPO) is one of the most popular deep reinforcement learning (RL) methods, achieving state-of-the-art performance across a wide range of challenging ta…