collaborators

5 papers

cs.LG2026

Offline Contextual Bandits in the Presence of New Actions

Ren Kishimoto, Tatsuhiro Shimizu, Kazuki Kawamura +6

Automated decision-making algorithms drive applications such as recommendation systems and search engines. These algorithms often rely on off-policy contextual bandits or off-polic…

cs.LG2026

Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies

Koichi Tanaka, Kazuki Kawamura, Takanori Muroi +6

Off-Policy Evaluation (OPE) is an important practical problem in algorithmic ranking systems, where the goal is to estimate the expected performance of a new ranking policy using o…

cs.AI2025

Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning

Haruka Kiyohara, Yusuke Narita, Yuta Saito +2

In many real recommender systems, novel items are added frequently over time. The importance of sufficiently presenting novel actions has widely been acknowledged for improving lon…

cs.IR2025

Counterfactual Reciprocal Recommender Systems for User-to-User Matching

Kazuki Kawamura, Takuma Udagawa, Kei Tateno

Reciprocal recommender systems (RRS) in dating, gaming, and talent platforms require mutual acceptance for a match. Logged data, however, over-represents popular profiles due to pa…

cs.IR2025

Not Just What, But When: Integrating Irregular Intervals to LLM for Sequential Recommendation

Wei-Wei Du, Takuma Udagawa, Kei Tateno

Time intervals between purchasing items are a crucial factor in sequential recommendation tasks, whereas existing approaches focus on item sequences and often overlook by assuming…