1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Yixiu Mao, Qi Wang, Yun Qu +2
Offline Reinforcement Learning (RL) suffers from the extrapolation error and value overestimation. From a generalization perspective, this issue can be attributed to the over-gener…