3 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.LG2011★ 3 cited
Reducing Commitment to Tasks with Off-Policy Hierarchical Reinforcement Learning
Mitchell Keith Bloch
In experimenting with off-policy temporal difference (TD) methods in hierarchical reinforcement learning (HRL) systems, we have observed unwanted on-policy learning under reproduci…
cs.LG2011★ 2 cited
Temporal Second Difference Traces
Mitchell Keith Bloch
Q-learning is a reliable but inefficient off-policy temporal-difference method, backing up reward only one step at a time. Replacing traces, using a recency heuristic, are more eff…