Recommendation Is a Dish Better Served Warm
arXiv:2508.07856 · doi:10.1145/3705328.3759331
Abstract
In modern recommender systems, experimental settings typically include filtering out cold users and items based on a minimum interaction threshold. However, these thresholds are often chosen arbitrarily and vary widely across studies, leading to inconsistencies that can significantly affect the comparability and reliability of evaluation results. In this paper, we systematically explore the cold-start boundary by examining the criteria used to determine whether a user or an item should be considered cold. Our experiments incrementally vary the number of interactions for different items during training, and gradually update the length of user interaction histories during inference. We investigate the thresholds across several widely used datasets, commonly represented in recent papers from top-tier conferences, and on multiple established recommender baselines. Our findings show that inconsistent selection of cold-start thresholds can either result in the unnecessary removal of valuable data or lead to the misclassification of cold instances as warm, introducing more noise into the system.
Accepted for ACM RecSys 2025. Author's version. The final published version will be available at the ACM Digital Library
References in corpus (9)
- Embarrassingly Shallow Autoencoders for Sparse Data
- Vista: A Visually, Socially, and Temporally-aware Model for Artistic Recommendation
- A Case Study on Sampling Strategies for Evaluating Neural Sequential Item Recommendation Models
- How Useful are Reviews for Recommendation? A Critical Review and Potential Improvements
- Turning Dross Into Gold Loss: is BERT4Rec really better than SASRec?
- Widespread Flaws in Offline Evaluation of Recommender Systems
- RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
- Dynamic Modeling of User Preferences for Stable Recommendations
- Maximum Impact with Fewer Features: Efficient Feature Selection for Cold-Start Recommenders through Collaborative Importance Weighting