3 papers
stat.ME2026
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
Kohsuke Kubota, Mitsuhiro Takahashi, Yuta Saito
Optimizing survival outcomes, such as patient survival or customer retention, is a critical objective in data-driven decision-making. Off-Policy Evaluation~(OPE) provides a powerfu…
cs.LG2025
MultiScale Contextual Bandits for Long Term Objectives
Richa Rastogi, Yuta Saito, Thorsten Joachims
The feedback that AI systems (e.g., recommender systems, chatbots) collect from user interactions is a crucial source of training data. While short-term feedback (e.g., clicks, eng…
cs.LG2024
Long-term Off-Policy Evaluation and Learning
Yuta Saito, Himan Abdollahpouri, Jesse Anderton +2
Short- and long-term outcomes of an algorithm often differ, with damaging downstream effects. A known example is a click-bait algorithm, which may increase short-term clicks but da…