1 paper
Botao Dong, Longyang Huang, Ning Pang +1
In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding th…