1 paper
Hao Hu, Yiqin Yang, Jianing Ye +7
Offline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and fu…