1 paper
Yuhui Wang, Miroslav Strupl, Francesco Faccio +5
Learning from multi-step off-policy data collected by a set of policies is a core problem of reinforcement learning (RL). Approaches based on importance sampling (IS) often suffer…