A Critical Study on Data Leakage in Recommender System Offline Evaluation
arXiv:2010.11060 · doi:10.1145/3569930
Abstract
Recommender models are hard to evaluate, particularly under offline setting. In this paper, we provide a comprehensive and critical analysis of the data leakage issue in recommender system offline evaluation. Data leakage is caused by not observing global timeline in evaluating recommenders, e.g., train/test data split does not follow global timeline. As a result, a model learns from the user-item interactions that are not expected to be available at prediction time. We first show the temporal dynamics of user-item interactions along global timeline, then explain why data leakage exists for collaborative filtering models. Through carefully designed experiments, we show that all models indeed recommend future items that are not available at the time point of a test instance, as the result of data leakage. The experiments are conducted with four widely used baseline models - BPR, NeuMF, SASRec, and LightGCN, on four popular offline datasets - MovieLens-25M, Yelp, Amazon-music, and Amazon-electronic, adopting leave-last-one-out data split. We further show that data leakage does impact models' recommendation accuracy. Their relative performance orders thus become unpredictable with different amount of leaked future data in training. To evaluate recommendation systems in a realistic manner in offline setting, we propose a timeline scheme, which calls for a revisit of the recommendation model design.
Accepted by TOIS
References in corpus (9)
- BPR: Bayesian Personalized Ranking from Implicit Feedback
- Session-based Social Recommendation via Dynamic Graph Attention Networks
- Causal Intervention for Leveraging Popularity Bias in Recommendation
- Elliot: a Comprehensive and Rigorous Framework for Reproducible Recommender Systems Evaluation
- GraphSAIL: Graph Structure Aware Incremental Learning for Recommender Systems
- How to Retrain Recommender System? A Sequential Meta-Learning Method
- A Re-visit of the Popularity Baseline in Recommender Systems
- Learning an Adaptive Meta Model-Generator for Incrementally Updating Recommender Systems
- Evaluation Metrics for Item Recommendation under Sampling
Cited by in corpus (17)
- Quality Metrics in Recommender Systems: Do We Calculate Metrics Consistently?
- Take a Fresh Look at Recommender Systems from an Evaluation Standpoint
- Does It Look Sequential? An Analysis of Datasets for Evaluation of Sequential Recommendations
- Revisiting BPR: A Replicability Study of a Common Recommender System Baseline
- Time to Split: Exploring Data Splitting Strategies for Offline Evaluation of Sequential Recommenders
- Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item Catalogs
- From Variability to Stability: Advancing RecSys Benchmarking Practices
- RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
- Understanding the Influence of Data Characteristics on the Performance of Point-of-Interest Recommendation Algorithms
- Impacts of Mainstream-Driven Algorithms on Recommendations for Children Across Domains: A Reproducibility Study
- Rolling Forward: Enhancing LightGCN with Causal Graph Convolution for Credit Bond Recommendation
- Beyond Collaborative Filtering: A Relook at Task Formulation in Recommender Systems
- Recommendation Is a Dish Better Served Warm
- eSASRec: Enhancing Transformer-based Recommendations in a Modular Fashion
- Correcting the LogQ Correction: Revisiting Sampled Softmax for Large-Scale Retrieval
- Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
- Point of Interest Recommendation: Pitfalls and Viable Solutions