Taking the Counterfactual Online: Efficient and Unbiased Online Evaluation for Ranking
arXiv:2007.12719 · doi:10.1145/3409256.3409820
Abstract
Counterfactual evaluation can estimate Click-Through-Rate (CTR) differences between ranking systems based on historical interaction data, while mitigating the effect of position bias and item-selection bias. We introduce the novel Logging-Policy Optimization Algorithm (LogOpt), which optimizes the policy for logging data so that the counterfactual estimate has minimal variance. As minimizing variance leads to faster convergence, LogOpt increases the data-efficiency of counterfactual estimation. LogOpt turns the counterfactual approach - which is indifferent to the logging policy - into an online approach, where the algorithm decides what rankings to display. We prove that, as an online evaluation method, LogOpt is unbiased w.r.t. position and item-selection bias, unlike existing interleaving methods. Furthermore, we perform large-scale experiments by simulating comparisons between thousands of rankers. Our results show that while interleaving methods make systematic errors, LogOpt is as efficient as interleaving without being biased.
ICTIR 2020
References in corpus (1)
Cited by in corpus (7)
- Unifying Online and Counterfactual Learning to Rank
- Reaching the End of Unbiasedness: Uncovering Implicit Limitations of Click-Based Learning to Rank
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization
- An Offline Metric for the Debiasedness of Click Models
- Learning from User Interactions with Rankings: A Unification of the Field
- Recent Advances in the Foundations and Applications of Unbiased Learning to Rank
- Exposure-Based Reinforcement Learning to Rank