1 paper
Austin A. Nguyen, Michael P. Wellman
Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectories. We consider the problem in a mixed-mo…