1 paper
Xuemin Hu, Shen Li, Yingfen Xu +2
Offline reinforcement learning (RL) can learn optimal policies from pre-collected offline datasets without interacting with the environment, but the sampled actions of the agent ca…