1 paper · 1 filter
Qingjun Wang, Hongtu Zhou, Hang Yu +5
Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods mitigate this issue by penalizing…