1 paper
Yi Yang, Zhennan Chen, Mingfeng Lv +3
Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be amplified by bootstrapping. Exi…