13 citations · 13 across the 1 of their papers we have counts for
1 paper
Zhizhou Ren, Guangxiang Zhu, Hao Hu +3
Double Q-learning is a classical method for reducing overestimation bias, which is caused by taking maximum estimated values in the Bellman operation. Its variants in the deep Q-le…