1 paper
Ke Jiang, Wen Jiang, Yao Li +1
We address the challenge of offline reinforcement learning using realistic data, specifically non-expert data collected through sub-optimal behavior policies. Under such circumstan…