1 paper
Omer Gottesman, Yao Liu, Scott Sussex +2
We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-pa…