1 paper
Julius Wagenbach, Matthia Sabatelli
We study whether the learning rate α, the discount factor γ and the reward signal r have an influence on the overestimation bias of the Q-Learning algorithm. Our preliminary…