Revisiting BPR: A Replicability Study of a Common Recommender System Baseline
arXiv:2409.14217 · doi:10.1145/3640457.3688073
Abstract
Bayesian Personalized Ranking (BPR), a collaborative filtering approach based on matrix factorization, frequently serves as a benchmark for recommender systems research. However, numerous studies often overlook the nuances of BPR implementation, claiming that it performs worse than newly proposed methods across various tasks. In this paper, we thoroughly examine the features of the BPR model, indicating their impact on its performance, and investigate open-source BPR implementations. Our analysis reveals inconsistencies between these implementations and the original BPR paper, leading to a significant decrease in performance of up to 50% for specific implementations. Furthermore, through extensive experiments on real-world datasets under modern evaluation settings, we demonstrate that with proper tuning of its hyperparameters, the BPR model can achieve performance levels close to state-of-the-art methods on the top-n recommendation tasks and even outperform them on specific datasets. Specifically, on the Million Song Dataset, the BPR model with hyperparameters tuning statistically significantly outperforms Mult-VAE by 10% in NDCG@100 with binary relevance function.
This paper is accepted at the Reproducibility track of the ACM RecSys '24 conference
References in corpus (12)
- Recurrent Neural Networks with Top-k Gains for Session-based Recommendations
- A Survey on Conversational Recommender Systems
- Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches
- Embarrassingly Shallow Autoencoders for Sparse Data
- Elliot: a Comprehensive and Rigorous Framework for Reproducible Recommender Systems Evaluation
- A Critical Study on Data Leakage in Recommender System Offline Evaluation
- A Re-visit of the Popularity Baseline in Recommender Systems
- Top-N Recommendation Algorithms: A Quest for the State-of-the-Art
- Widespread Flaws in Offline Evaluation of Recommender Systems
- Debiased Explainable Pairwise Ranking from Implicit Feedback
- The Effect of Third Party Implementations on Reproducibility
- From Variability to Stability: Advancing RecSys Benchmarking Practices