Policy-Aware Unbiased Learning to Rank for Top-k Rankings
arXiv:2005.09035 · doi:10.1145/3397271.3401102
Abstract
Counterfactual Learning to Rank (LTR) methods optimize ranking systems using logged user interactions that contain interaction biases. Existing methods are only unbiased if users are presented with all relevant items in every ranking. There is currently no existing counterfactual unbiased LTR method for top-k rankings. We introduce a novel policy-aware counterfactual estimator for LTR metrics that can account for the effect of a stochastic logging policy. We prove that the policy-aware estimator is unbiased if every relevant item has a non-zero probability to appear in the top-k ranking. Our experimental results show that the performance of our estimator is not affected by the size of k: for any k, the policy-aware estimator reaches the same retrieval performance while learning from top-k feedback as when learning from feedback on the full ranking. Lastly, we introduce novel extensions of traditional LTR methods to perform counterfactual LTR and to optimize top-k metrics. Together, our contributions introduce the first policy-aware unbiased LTR approach that learns from top-k feedback and optimizes top-k metrics. As a result, counterfactual LTR is now applicable to the very prevalent top-k ranking setting in search and recommendation.
SIGIR 2020 full conference paper
References in corpus (2)
Cited by in corpus (21)
- When Inverse Propensity Scoring does not Work: Affine Corrections for Unbiased Learning to Rank
- Deconfounded Causal Collaborative Filtering
- Causal Collaborative Filtering
- Doubly-Robust Estimation for Correcting Position-Bias in Click Feedback for Unbiased Learning to Rank
- Reaching the End of Unbiasedness: Uncovering Implicit Limitations of Click-Based Learning to Rank
- Taking the Counterfactual Online: Efficient and Unbiased Online Evaluation for Ranking
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization
- A Probabilistic Position Bias Model for Short-Video Recommendation Feeds
- Mixture-Based Correction for Position and Trust Bias in Counterfactual Learning to Rank
- The Bandwagon Effect: Not Just Another Bias
- An Offline Metric for the Debiasedness of Click Models
- Adaptive Orchestration of Modular Generative Information Access Systems
- Robust Generalization and Safe Query-Specialization in Counterfactual Learning to Rank
- Marginal-Certainty-aware Fair Ranking Algorithm
- On the Impact of Outlier Bias on User Clicks
- Recent Advances in the Foundations and Applications of Unbiased Learning to Rank
- Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank
- FARA: Future-aware Ranking Algorithm for Fairness Optimization
- Learning to Rank with Variable Result Presentation Lengths
- Investigating the Robustness of Counterfactual Learning to Rank Models: A Reproducibility Study
- Towards Two-Stage Counterfactual Learning to Rank