1 paper
Amna Najib, Stefan Depeweg, Phillip Swazinna
Batch reinforcement learning enables policy learning without direct interaction with the environment during training, relying exclusively on previously collected sets of interactio…