38 citations · 116 across the 12 of their papers we have counts for
3 papers · 1 filter
Data-Driven Off-Policy Estimator Selection: An Application in User Marketing on An Online Content Delivery Service
Yuta Saito, Takuma Udagawa, Kei Tateno
Off-policy evaluation (OPE) is the method that attempts to estimate the performance of decision making policies using historical data generated by different policies without conduc…
Accelerating Offline Reinforcement Learning Application in Real-Time Bidding and Recommendation: Potential Use of Simulation
Haruka Kiyohara, Kosuke Kawakami, Yuta Saito
In recommender systems (RecSys) and real-time bidding (RTB) for online advertisements, we often try to optimize sequential decision making using bandit and reinforcement learning (…
Evaluating the Robustness of Off-Policy Evaluation
Yuta Saito, Takuma Udagawa, Haruka Kiyohara +3
Off-policy Evaluation (OPE), or offline evaluation in general, evaluates the performance of hypothetical policies leveraging only offline log data. It is particularly useful in app…