4 citations · 4 across the 1 of their papers we have counts for
1 paper
Sabrina Hoppe, Marc Toussaint
In state of the art model-free off-policy deep reinforcement learning, a replay memory is used to store past experience and derive all network updates. Even if both state and actio…