On the Difficulty of Evaluating Baselines: A Study on Recommender Systems
arXiv:1905.01395
Abstract
Numerical evaluations with comparisons to baselines play a central role when judging research in recommender systems. In this paper, we show that running baselines properly is difficult. We demonstrate this issue on two extensively studied datasets. First, we show that results for baselines that have been used in numerous publications over the past five years for the Movielens 10M benchmark are suboptimal. With a careful setup of a vanilla matrix factorization baseline, we are not only able to improve upon the reported results for this baseline but even outperform the reported results of any newly proposed method. Secondly, we recap the tremendous effort that was required by the community to obtain high quality results for simple methods on the Netflix Prize. Our results indicate that empirical findings in research papers are questionable unless they were obtained on standardized benchmarks where baselines have been tuned extensively by the research community.
Cited by in corpus (18)
- Quality Metrics in Recommender Systems: Do We Calculate Metrics Consistently?
- On Sampling Top-K Recommendation Evaluation
- Towards a Better Understanding of Linear Models for Recommendation
- Critically Examining the Claimed Value of Convolutions over User-Item Embedding Maps for Recommender Systems
- Do Offline Metrics Predict Online Performance in Recommender Systems?
- A Lightweight Method for Modeling Confidence in Recommendations with Learned Beta Distributions
- Initialization Matters: Regularizing Manifold-informed Initialization for Neural Recommendation Systems
- Exploring Data Splitting Strategies for the Evaluation of Recommendation Models
- Content Based Player and Game Interaction Model for Game Recommendation in the Cold Start setting
- It's Enough: Relaxing Diagonal Constraints in Linear Autoencoders for Recommendation
- On Estimating Recommendation Evaluation Metrics under Sampling
- UserReg: A Simple but Strong Model for Rating Prediction
- Review Regularized Neural Collaborative Filtering
- Prevention is Better than Cure: Handling Basis Collapse and Transparency in Dense Networks
- Local Search Algorithms for Rank-Constrained Convex Optimization
- New Recommendation Algorithm for Implicit Data Motivated by the Multivariate Normal Distribution
- Statistical Inference: The Missing Piece of RecSys Experiment Reliability Discourse
- Ideas for Improving the Field of Machine Learning: Summarizing Discussion from the NeurIPS 2019 Retrospectives Workshop