14 citations · 25 across the 4 of their papers we have counts for
1 paper · 1 filter
Chenjun Xiao, Han Wang, Yangchen Pan +2
Reinforcement learning (RL) agents can leverage batches of previously collected data to extract a reasonable control policy. An emerging issue in this offline RL setting, however,…