45 citations · 246 across the 39 of their papers we have counts for
1 paper · 2 filters
Aishwarya Mandyam, Jason Meng, Ge Gao +4
Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary dataset…