2 papers
stat.ML2025
A Principled Path to Fitted Distributional Evaluation
Sungee Hong, Jiayi Wang, Zhengling Qi +1
In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a differen…
stat.ML2025
Distributional Off-policy Evaluation with Bellman Residual Minimization
Sungee Hong, Zhengling Qi, Raymond K. W. Wong
We study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different po…