1 paper
Xiao-Yin Liu, Guotao Li, Xiao-Hu Zhou +1
Offline preference-based reinforcement learning (PbRL) provides an effective way to overcome the challenges of designing reward and the high costs of online interaction. However, s…