Policy-Gradient Training of Fair and Unbiased Ranking Functions
arXiv:1911.08054 · doi:10.1145/3404835.3462953
Abstract
While implicit feedback (e.g., clicks, dwell times, etc.) is an abundant and attractive source of data for learning to rank, it can produce unfair ranking policies for both exogenous and endogenous reasons. Exogenous reasons typically manifest themselves as biases in the training data, which then get reflected in the learned ranking policy and often lead to rich-get-richer dynamics. Moreover, even after the correction of such biases, reasons endogenous to the design of the learning algorithm can still lead to ranking policies that do not allocate exposure among items in a fair way. To address both exogenous and endogenous sources of unfairness, we present the first learning-to-rank approach that addresses both presentation bias and merit-based fairness of exposure simultaneously. Specifically, we define a class of amortized fairness-of-exposure constraints that can be chosen based on the needs of an application, and we show how these fairness criteria can be enforced despite the selection biases in implicit feedback data. The key result is an efficient and flexible policy-gradient algorithm, called FULTR, which is the first to enable the use of counterfactual estimators for both utility estimation and fairness constraints. Beyond the theoretical justification of the framework, we show empirically that the proposed algorithm can learn accurate and fair ranking policies from biased and noisy feedback.
References in corpus (6)
- Fairness of Exposure in Rankings
- Fairness-Aware Ranking in Search & Recommendation Systems with Application to LinkedIn Talent Search
- Equity of Attention: Amortizing Individual Fairness in Rankings
- Controlling Fairness and Bias in Dynamic Learning-to-Rank
- Estimating Position Bias without Intrusive Interventions
- Intervention Harvesting for Context-Dependent Examination-Bias Estimation
Cited by in corpus (10)
- Fairness in Recommender Systems: Research Landscape and Future Directions
- Fair Ranking as Fair Division: Impact-Based Individual Fairness in Ranking
- Fairness of Exposure in Light of Incomplete Exposure Estimation
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization
- Probabilistic Permutation Graph Search: Black-Box Optimization for Fairness in Ranking
- Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank
- Recent Advances in the Foundations and Applications of Unbiased Learning to Rank
- Scalable and Provably Fair Exposure Control for Large-Scale Recommender Systems
- Towards Two-Stage Counterfactual Learning to Rank
- Learning to Re-rank with Constrained Meta-Optimal Transport