Time to Split: Exploring Data Splitting Strategies for Offline Evaluation of Sequential Recommenders
arXiv:2507.16289 · doi:10.1145/3705328.3748164
Abstract
Modern sequential recommender systems, ranging from lightweight transformer-based variants to large language models, have become increasingly prominent in academia and industry due to their strong performance in the next-item prediction task. Yet common evaluation protocols for sequential recommendations remain insufficiently developed: they often fail to reflect the corresponding recommendation task accurately, or are not aligned with real-world scenarios. Although the widely used leave-one-out split matches next-item prediction, it permits the overlap between training and test periods, which leads to temporal leakage and unrealistically long test horizon, limiting real-world relevance. Global temporal splitting addresses these issues by evaluating on distinct future periods. However, its applications to sequential recommendations remain loosely defined, particularly in terms of selecting target interactions and constructing a validation subset that provides necessary consistency between validation and test metrics. In this paper, we demonstrate that evaluation outcomes can vary significantly across splitting strategies, influencing model rankings and practical deployment decisions. To improve reproducibility in both academic and industrial settings, we systematically compare different splitting strategies for sequential recommendations across multiple datasets and established baselines. Our findings show that prevalent splits, such as leave-one-out, may be insufficiently aligned with more realistic evaluation strategies. Code: https://github.com/monkey0head/time-to-split
Accepted for ACM RecSys 2025. Author's version. The final published version will be available at the ACM Digital Library
References in corpus (22)
- Contrastive Learning for Representation Degeneration Problem in Sequential Recommendation
- Translation-based Recommendation
- Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches
- Evaluation of Session-based Recommendation Algorithms
- Leveraging Large Language Models for Sequential Recommendation
- Elliot: a Comprehensive and Rigorous Framework for Reproducible Recommender Systems Evaluation
- A Critical Study on Data Leakage in Recommender System Offline Evaluation
- Denoising Self-attentive Sequential Recommendation
- A Case Study on Sampling Strategies for Evaluating Neural Sequential Item Recommendation Models
- How Useful are Reviews for Recommendation? A Critical Review and Potential Improvements
- Turning Dross Into Gold Loss: is BERT4Rec really better than SASRec?
- gSASRec: Reducing Overconfidence in Sequential Recommendation Trained with Negative Sampling
- Take a Fresh Look at Recommender Systems from an Evaluation Standpoint
- Widespread Flaws in Offline Evaluation of Recommender Systems
- DaisyRec 2.0: Benchmarking Recommendation for Rigorous Evaluation
- Does It Look Sequential? An Analysis of Datasets for Evaluation of Sequential Recommendations
- Microsoft Recommenders: Tools to Accelerate Developing Recommender Systems
- The Effect of Third Party Implementations on Reproducibility
- Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item Catalogs
- From Variability to Stability: Advancing RecSys Benchmarking Practices
- RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
- RePlay: a Recommendation Framework for Experimentation and Production Use