1 citations · 1 across the 3 of their papers we have counts for
1 paper · 2 filters
Leonard S. Pleiss, James Harrison, Maximilian Schiffer
Many value-based deep reinforcement learning algorithms rely on target networks - lagged copies of the online network - to stabilize training. While effective, this mechanism intro…