1 paper
Yang Hu, Tianyi Chen, Na Li +2
Off-policy evaluation (OPE) is one of the most fundamental problems in reinforcement learning (RL) to estimate the expected long-term payoff of a given target policy with only expe…