4 papers
Offline Contextual Bandits in the Presence of New Actions
Ren Kishimoto, Tatsuhiro Shimizu, Kazuki Kawamura +6
Automated decision-making algorithms drive applications such as recommendation systems and search engines. These algorithms often rely on off-policy contextual bandits or off-polic…
Off-Policy Learning with Limited Supply
Koichi Tanaka, Ren Kishimoto, Bushun Kawagishi +4
We study off-policy learning (OPL) in contextual bandits, which plays a key role in a wide range of real-world applications such as recommendation systems and online advertising. T…
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
Koichi Tanaka, Kazuki Kawamura, Takanori Muroi +6
Off-Policy Evaluation (OPE) is an important practical problem in algorithmic ranking systems, where the goal is to estimate the expected performance of a new ranking policy using o…
Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning
Haruka Kiyohara, Yusuke Narita, Yuta Saito +2
In many real recommender systems, novel items are added frequently over time. The importance of sufficiently presenting novel actions has widely been acknowledged for improving lon…