Showing 2024 · cs.LGShow all
2 papers · 2 filters
cs.LG2024
Demystifying the Recency Heuristic in Temporal-Difference Learning
Brett Daley, Marlos C. Machado, Martha White
The recency heuristic in reinforcement learning is the assumption that stimuli that occurred closer in time to an acquired reward should be more heavily reinforced. The recency heu…
cs.LG2024
Averaging -step Returns Reduces Variance in Reinforcement Learning
Brett Daley, Martha White, Marlos C. Machado
Multistep returns, such as -step returns and -returns, are commonly used to improve the sample efficiency of reinforcement learning (RL) methods. The variance of the multiste…