1 paper
Hyunwoo Kim, Hyo Kyung Lee
Off-policy reinforcement learning suffers from extrapolation errors when a learned policy selects actions that are weakly supported in the replay buffer. In this study, we address…