18 citations · 48 across the 12 of their papers we have counts for
1 paper · 1 filter
Aurélien F. Bibaut, Ivana Malenica, Nikos Vlassis +1
We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have…