17 citations · 28 across the 8 of their papers we have counts for
1 paper · 1 filter
Aurélien F. Bibaut, Ivana Malenica, Nikos Vlassis +1
We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have…