11 citations · 17 across the 13 of their papers we have counts for
1 paper · 2 filters
Armaan A. Abraham, Lucy Xiaoyang Shi, Chelsea Finn
Off-policy, value-based reinforcement learning methods such as Q-learning are appealing because they can learn from arbitrary experience, including data collected by older policies…