1 paper
Matthew Thomas Jackson, Michael Tryfan Matthews, Cong Lu +3
In many real-world settings, agents must learn from an offline dataset gathered by some prior behavior policy. Such a setting naturally leads to distribution shift between the beha…