356 citations · 2k across the 65 of their papers we have counts for
1 paper · 1 filter
Jianlan Luo, Perry Dong, Jeffrey Wu +3
The offline reinforcement learning (RL) paradigm provides a general recipe to convert static behavior datasets into policies that can perform better than the policy that collected…