2 citations · 3 across the 4 of their papers we have counts for
1 paper · 1 filter
Yixiu Mao, Hongchang Zhang, Chen Chen +2
Offline reinforcement learning suffers from the out-of-distribution issue and extrapolation error. Most policy constraint methods regularize the density of the trained policy towar…