1 paper
Zhengbang Zhu, Minghuan Liu, Liyuan Mao +5
Offline reinforcement learning (RL) aims to learn policies from pre-existing datasets without further interactions, making it a challenging task. Q-learning algorithms struggle wit…