1 paper · 1 filter
Sungee Hong, Jiayi Wang, Zhengling Qi +1
In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a differen…