1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 1 cited
Improving the Efficiency of Off-Policy Reinforcement Learning by Accounting for Past Decisions
Brett Daley, Christopher Amato
Off-policy learning from multistep returns is crucial for sample-efficient reinforcement learning, particularly in the experience replay setting now commonly used with deep neural…
cs.LG2021
Virtual Replay Cache
Brett Daley, Christopher Amato
Return caching is a recent strategy that enables efficient minibatch training with multistep estimators (e.g. the λ-return) for deep reinforcement learning. By precomputing return…