1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Armaan A. Abraham, Lucy Xiaoyang Shi, Chelsea Finn
Off-policy, value-based reinforcement learning methods such as Q-learning are appealing because they can learn from arbitrary experience, including data collected by older policies…