Take a Fresh Look at Recommender Systems from an Evaluation Standpoint
arXiv:2210.04149 · doi:10.1145/3539618.3591931
Abstract
Recommendation has become a prominent area of research in the field of Information Retrieval (IR). Evaluation is also a traditional research topic in this community. Motivated by a few counter-intuitive observations reported in recent studies, this perspectives paper takes a fresh look at recommender systems from an evaluation standpoint. Rather than examining metrics like recall, hit rate, or NDCG, or perspectives like novelty and diversity, the key focus here is on how these metrics are calculated when evaluating a recommender algorithm. Specifically, the commonly used train/test data splits and their consequences are re-examined. We begin by examining common data splitting methods, such as random split or leave-one-out, and discuss why the popularity baseline is poorly defined under such splits. We then move on to explore the two implications of neglecting a global timeline during evaluation: data leakage and oversimplification of user preference modeling. Afterwards, we present new perspectives on recommender systems, including techniques for evaluating algorithm performance that more accurately reflect real-world scenarios, and possible approaches to consider decision contexts in user preference modeling.
Accepted for SIGIR 2023 (Perspectives Paper Track)
References in corpus (4)
- Quality Metrics in Recommender Systems: Do We Calculate Metrics Consistently?
- A Re-visit of the Popularity Baseline in Recommender Systems
- FINN.no Slates Dataset: A new Sequential Dataset Logging Interactions, allViewed Items and Click Responses/No-Click for Recommender Systems Research
- Do Loyal Users Enjoy Better Recommendations? Understanding Recommender Accuracy from a Time Perspective
Cited by in corpus (9)
- Does It Look Sequential? An Analysis of Datasets for Evaluation of Sequential Recommendations
- Time to Split: Exploring Data Splitting Strategies for Offline Evaluation of Sequential Recommenders
- From Variability to Stability: Advancing RecSys Benchmarking Practices
- Rolling Forward: Enhancing LightGCN with Causal Graph Convolution for Credit Bond Recommendation
- Beyond Collaborative Filtering: A Relook at Task Formulation in Recommender Systems
- Recommendation Is a Dish Better Served Warm
- Impression-Aware Recommender Systems
- Correcting the LogQ Correction: Revisiting Sampled Softmax for Large-Scale Retrieval
- Extending MovieLens-32M to Provide New Evaluation Objectives