1 paper
Yaomin Wang, Jianting Pan, Ran Tian +4
The discount factor in reinforcement learning controls both the effective planning horizon and the strength of bootstrapping, yet most deep RL methods use a single fixed value acro…