157 citations · 546 across the 9 of their papers we have counts for
1 paper · 1 filter
Dillon Sandhu, Ronald Parr
We revisit a classic "chicken-and-egg" problem in reinforcement learning: to safely improve a policy, the value function must be accurate on the state-visitation distribution of th…