2 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Dhruva Tirumala, Thomas Lampe, Jose Enrique Chen +9
Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to exte…