38 citations · 73 across the 8 of their papers we have counts for
8 papers
Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking
Haruka Kiyohara, Mihaela Curmei, Ariel Evnine +7
Large-scale search, recommendation, and retrieval-augmented generation (RAG) systems typically employ a two-stage architecture: an early-stage ranker (ESR) generates a candidate se…
Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning
Haruka Kiyohara, Yusuke Narita, Yuta Saito +2
In many real recommender systems, novel items are added frequently over time. The importance of sufficiently presenting novel actions has widely been acknowledged for improving lon…
Prompt Optimization with Logged Bandit Data
Haruka Kiyohara, Daniel Yiming Cao, Yuta Saito +1
We study how to use naturally available user feedback, such as clicks, to optimize large language model (LLM) pipelines for generating personalized sentences using prompts. Naive a…
Policy Design for Two-sided Platforms with Participation Dynamics
Haruka Kiyohara, Fan Yao, Sarah Dean
In two-sided platforms (e.g., video streaming or e-commerce), viewers and providers engage in interactive dynamics: viewers benefit from increases in provider populations, while pr…
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
Tatsuhiro Shimizu, Koichi Tanaka, Ren Kishimoto +3
We explore off-policy evaluation and learning (OPE/L) in contextual combinatorial bandits (CCB), where a policy selects a subset in the action space. For example, it might choose a…
Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model
Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro +3
In real-world recommender systems and search engines, optimizing ranking decisions to present a ranked list of relevant items is critical. Off-policy evaluation (OPE) for ranking p…