5 papers
Offline Contextual Bandits in the Presence of New Actions
Ren Kishimoto, Tatsuhiro Shimizu, Kazuki Kawamura +6
Automated decision-making algorithms drive applications such as recommendation systems and search engines. These algorithms often rely on off-policy contextual bandits or off-polic…
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows
Kazuki Kawamura, Satoshi Waki, Kei Tateno
Multi-agent LLM workflows -- systems composed of multiple role-specific LLM calls -- often outperform single-prompt baselines, but they remain difficult to debug and refine. Failur…
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
Koichi Tanaka, Kazuki Kawamura, Takanori Muroi +6
Off-Policy Evaluation (OPE) is an important practical problem in algorithmic ranking systems, where the goal is to estimate the expected performance of a new ranking policy using o…
Counterfactual Reciprocal Recommender Systems for User-to-User Matching
Kazuki Kawamura, Takuma Udagawa, Kei Tateno
Reciprocal recommender systems (RRS) in dating, gaming, and talent platforms require mutual acceptance for a match. Logged data, however, over-represents popular profiles due to pa…
Parallel and Mini-Batch Stable Matching for Large-Scale Reciprocal Recommender Systems
Kento Nakada, Kazuki Kawamura, Ryosuke Furukawa
Reciprocal recommender systems (RRSs) are crucial in online two-sided matching platforms, such as online job or dating markets, as they need to consider the preferences of both sid…