1 paper
Catalin E. Brita, Stephan Bongers, Frans A. Oliehoek
In offline reinforcement learning, deriving an effective policy from a pre-collected set of experiences is challenging due to the distribution mismatch between the target policy an…