Distributionally-Informed Recommender System Evaluation
arXiv:2309.05892 · doi:10.1145/3613455
Abstract
Current practice for evaluating recommender systems typically focuses on point estimates of user-oriented effectiveness metrics or business metrics, sometimes combined with additional metrics for considerations such as diversity and novelty. In this paper, we argue for the need for researchers and practitioners to attend more closely to various distributions that arise from a recommender system (or other information access system) and the sources of uncertainty that lead to these distributions. One immediate implication of our argument is that both researchers and practitioners must report and examine more thoroughly the distribution of utility between and within different stakeholder groups. However, distributions of various forms arise in many more aspects of the recommender systems experimental process, and distributional thinking has substantial ramifications for how we design, evaluate, and present recommender systems evaluation and research results. Leveraging and emphasizing distributions in the evaluation of recommender systems is a necessary step to ensure that the systems provide appropriate and equitably-distributed benefit to the people they affect.
Accepted to ACM Transactions on Recommender Systems
References in corpus (4)
- Auditing Search Engines for Differential Satisfaction Across Demographics
- Statistical Significance Testing in Information Retrieval: An Empirical Analysis of Type I, Type II and Type III Errors
- Using Score Distributions to Compare Statistical Significance Tests for Information Retrieval Evaluation
- Estimating Error and Bias in Offline Evaluation Results