1 paper
Minting Pan, Yitao Zheng, Jiajian Li +2
Offline reinforcement learning (RL) enables policy optimization using static datasets, avoiding the risks and costs of extensive real-world exploration. However, it struggles with…