5 papers
Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior
Kosuke Iguchi, Ren Kishimoto
Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combi…
Offline Contextual Bandits in the Presence of New Actions
Ren Kishimoto, Tatsuhiro Shimizu, Kazuki Kawamura +6
Automated decision-making algorithms drive applications such as recommendation systems and search engines. These algorithms often rely on off-policy contextual bandits or off-polic…
Off-Policy Learning with Limited Supply
Koichi Tanaka, Ren Kishimoto, Bushun Kawagishi +4
We study off-policy learning (OPL) in contextual bandits, which plays a key role in a wide range of real-world applications such as recommendation systems and online advertising. T…
Beyond Match Maximization and Fairness: Retention-Optimized Two-Sided Matching
Ren Kishimoto, Rikiya Takehi, Koichi Tanaka +4
On two-sided matching platforms such as online dating and recruiting, recommendation algorithms often aim to maximize the total number of matches. However, this objective creates a…
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
Tatsuhiro Shimizu, Koichi Tanaka, Ren Kishimoto +3
We explore off-policy evaluation and learning (OPE/L) in contextual combinatorial bandits (CCB), where a policy selects a subset in the action space. For example, it might choose a…