2 papers
cs.LG2022
Robust Action Gap Increasing with Clipped Advantage Learning
Zhe Zhang, Yaozhong Gan, Xiaoyang Tan
Advantage Learning (AL) seeks to increase the action gap between the optimal action and its competitors, so as to improve the robustness to estimation errors. However, the method b…
cs.LG2022
Smoothing Advantage Learning
Yaozhong Gan, Zhe Zhang, Xiaoyang Tan
Advantage learning (AL) aims to improve the robustness of value-based reinforcement learning against estimation errors with action-gap-based regularization. Unfortunately, the meth…