69 citations · 144 across the 5 of their papers we have counts for
1 paper · 1 filter
Melrose Roderick, Gaurav Manek, Felix Berkenkamp +1
A key problem in off-policy Reinforcement Learning (RL) is the mismatch, or distribution shift, between the dataset and the distribution over states and actions visited by the lear…