Doubly-Robust Estimation for Correcting Position-Bias in Click Feedback for Unbiased Learning to Rank
arXiv:2203.17118 · doi:10.1145/3569453
Abstract
Clicks on rankings suffer from position-bias: generally items on lower ranks are less likely to be examined - and thus clicked - by users, in spite of their actual preferences between items. The prevalent approach to unbiased click-based learning-to-rank (LTR) is based on counterfactual inverse-propensity-scoring (IPS) estimation. In contrast with general reinforcement learning, counterfactual doubly-robust (DR) estimation has not been applied to click-based LTR in previous literature. In this paper, we introduce a novel DR estimator that is the first DR approach specifically designed for position-bias. The difficulty with position-bias is that the treatment - user examination - is not directly observable in click data. As a solution, our estimator uses the expected treatment per rank, instead of the actual treatment that existing DR estimators use. Our novel DR estimator has more robust unbiasedness conditions than the existing IPS approach, and in addition, provides enormous decreases in variance: our experimental results indicate it requires several orders of magnitude fewer datapoints to converge at optimal performance. For the unbiased LTR field, our DR estimator contributes both increases in state-of-the-art performance and the most robust theoretical guarantees of all known LTR estimators.
References in corpus (9)
- Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data
- Doubly Robust Policy Evaluation and Optimization
- Correcting for Selection Bias in Learning-to-rank Systems
- Unifying Online and Counterfactual Learning to Rank
- Policy-Aware Unbiased Learning to Rank for Top-k Rankings
- Cascade Model-based Propensity Estimation for Counterfactual Learning to Rank
- Non-Clicks Mean Irrelevant? Propensity Ratio Scoring As a Correction
- Unbiased Learning to Rank: Online or Offline?
- Robust Generalization and Safe Query-Specialization in Counterfactual Learning to Rank
Cited by in corpus (8)
- A Probabilistic Position Bias Model for Short-Video Recommendation Feeds
- Going Beyond Popularity and Positivity Bias: Correcting for Multifactorial Bias in Recommender Systems
- Recent Advances in the Foundations and Applications of Unbiased Learning to Rank
- Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank
- Meta Learning to Rank for Sparsely Supervised Queries
- Towards Two-Stage Counterfactual Learning to Rank
- An Epistemic Position-Based Click Model: From Interactions to Epistemic Distributions of Relevance and Bias
- Exposure-Based Reinforcement Learning to Rank