2 papers
cs.LG2022
Counterfactual Learning with General Data-generating Policies
Yusuke Narita, Kyohei Okumura, Akihiro Shimizu +1
Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE…
cs.LG2018
Efficient Counterfactual Learning from Bandit Feedback
Yusuke Narita, Shota Yasui, Kohei Yata
What is the most statistically efficient way to do off-policy evaluation and optimization with batch data from bandit feedback? For log data generated by contextual bandit algorith…