38 citations · 87 across the 14 of their papers we have counts for
5 papers · 1 filter
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
Tatsuhiro Shimizu, Koichi Tanaka, Ren Kishimoto +3
We explore off-policy evaluation and learning (OPE/L) in contextual combinatorial bandits (CCB), where a policy selects a subset in the action space. For example, it might choose a…
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
Haruka Kiyohara, Masahiro Nomura, Yuta Saito
We study off-policy evaluation (OPE) in the problem of slate contextual bandits where a policy selects multi-dimensional actions known as slates. This problem is widespread in reco…
Off-Policy Evaluation of Ranking Policies under Diverse User Behavior
Haruka Kiyohara, Masatoshi Uehara, Yusuke Narita +3
Ranking interfaces are everywhere in online platforms. There is thus an ever growing interest in their Off-Policy Evaluation (OPE), aiming towards an accurate performance evaluatio…
Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model
Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro +3
In real-world recommender systems and search engines, optimizing ranking decisions to present a ranked list of relevant items is critical. Off-policy evaluation (OPE) for ranking p…
Evaluating the Robustness of Off-Policy Evaluation
Yuta Saito, Takuma Udagawa, Haruka Kiyohara +3
Off-policy Evaluation (OPE), or offline evaluation in general, evaluates the performance of hypothetical policies leveraging only offline log data. It is particularly useful in app…