1 paper
Pu Li, Tao Tan, Hong Xie +2
This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. We find…