1 paper
Anthony GX-Chen, Veronica Chelu, Blake A. Richards +1
Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value func…