1 paper
Tenglong Liu, Yang Li, Yixing Lan +3
In offline reinforcement learning, the challenge of out-of-distribution (OOD) is pronounced. To address this, existing methods often constrain the learned policy through policy reg…