1 paper
Shirong Liu, Chenjia Bai, Zixian Guo +3
Policy constraint methods in offline reinforcement learning employ additional regularization techniques to constrain the discrepancy between the learned policy and the offline data…