1 paper
Lu Guo, Yixiang Shan, Zhengbang Zhu +5
Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions. However, its reliance on finite static datasets…