24 citations · 24 across the 1 of their papers we have counts for
1 paper
Chenjia Bai, Lingxiao Wang, Zhuoran Yang +4
Offline Reinforcement Learning (RL) aims to learn policies from previously collected datasets without exploring the environment. Directly applying off-policy algorithms to offline…