activity
20192022
most citedDoubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model

38 citations · 74 across the 6 of their papers we have counts for

collaborators

9 papers

cs.IR20224 cited

A Real-World Implementation of Unbiased Lift-based Bidding System

Daisuke Moriwaki, Yuta Hayakawa, Akira Matsui +3

In display ad auctions of Real-Time Bid-ding (RTB), a typical Demand-Side Platform (DSP)bids based on the predicted probability of click and conversion right after an ad impression…

stat.ML202238 cited

Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model

Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro +3

In real-world recommender systems and search engines, optimizing ranking decisions to present a ranked list of relevant items is critical. Off-policy evaluation (OPE) for ranking p…

cs.AI20211 cited

Data-Driven Off-Policy Estimator Selection: An Application in User Marketing on An Online Content Delivery Service

Yuta Saito, Takuma Udagawa, Kei Tateno

Off-policy evaluation (OPE) is the method that attempts to estimate the performance of decision making policies using historical data generated by different policies without conduc…

cs.LG20212 cited

Accelerating Offline Reinforcement Learning Application in Real-Time Bidding and Recommendation: Potential Use of Simulation

Haruka Kiyohara, Kosuke Kawakami, Yuta Saito

In recommender systems (RecSys) and real-time bidding (RTB) for online advertisements, we often try to optimize sequential decision making using bandit and reinforcement learning (…

stat.ML202125 cited

Evaluating the Robustness of Off-Policy Evaluation

Yuta Saito, Takuma Udagawa, Haruka Kiyohara +3

Off-policy Evaluation (OPE), or offline evaluation in general, evaluates the performance of hypothetical policies leveraging only offline log data. It is particularly useful in app…

cs.LG2020

Optimal Off-Policy Evaluation from Multiple Logging Policies

Nathan Kallus, Yuta Saito, Masatoshi Uehara

We study off-policy evaluation (OPE) from multiple logging policies, each generating a dataset of fixed size, i.e., stratified sampling. Previous work noted that in this setting th…