1 paper
Yifei Chen, Shaoqin Zhu, Xiaoqiang Ji
Offline reinforcement learning methods typically enforce strict constraints to ensure safety; yet this rigidity often prevents the discovery of optimal behaviors outside the immedi…