1 paper
Yuanhao Chen, Qi Liu, Pengbin Chen +2
Offline reinforcement learning (RL) aims to learn a policy that maximizes the expected return using a given static dataset of transitions. However, offline RL faces the distributio…