DaisyRec 2.0: Benchmarking Recommendation for Rigorous Evaluation
arXiv:2206.10848 · doi:10.1109/TPAMI.2022.3231891
Abstract
Recently, one critical issue looms large in the field of recommender systems -- there are no effective benchmarks for rigorous evaluation -- which consequently leads to unreproducible evaluation and unfair comparison. We, therefore, conduct studies from the perspectives of practical theory and experiments, aiming at benchmarking recommendation for rigorous evaluation. Regarding the theoretical study, a series of hyper-factors affecting recommendation performance throughout the whole evaluation chain are systematically summarized and analyzed via an exhaustive review on 141 papers published at eight top-tier conferences within 2017-2020. We then classify them into model-independent and model-dependent hyper-factors, and different modes of rigorous evaluation are defined and discussed in-depth accordingly. For the experimental study, we release DaisyRec 2.0 library by integrating these hyper-factors to perform rigorous evaluation, whereby a holistic empirical study is conducted to unveil the impacts of different hyper-factors on recommendation performance. Supported by the theoretical and experimental studies, we finally create benchmarks for rigorous evaluation by proposing standardized procedures and providing performance of ten state-of-the-arts across six evaluation metrics on six datasets as a reference for later study. Overall, our work sheds light on the issues in recommendation evaluation, provides potential solutions for rigorous evaluation, and lays foundation for further investigation.
DaisyRec-v2.0 contains a Python toolkit developed for benchmarking top-N recommendation task. Our code is available at https://github.com/recsys-benchmark/DaisyRec-v2.0
References in corpus (18)
- KGAT: Knowledge Graph Attention Network for Recommendation
- Knowledge Graph Convolutional Networks for Recommender Systems
- Robust Real-World Image Super-Resolution against Adversarial Attacks
- Unifying Knowledge Graph Learning and Recommendation: Towards a Better Understanding of User Preferences
- Disentangled Graph Collaborative Filtering
- RecVAE: a New Variational Autoencoder for Top-N Recommendations with Implicit Feedback
- DisenHAN: Disentangled Heterogeneous Graph Attention Network for Recommendation
- Neural Interactive Collaborative Filtering
- On the Difficulty of Evaluating Baselines: A Study on Recommender Systems
- CAFE: Coarse-to-Fine Neural Symbolic Reasoning for Explainable Recommendation
- DE-RRD: A Knowledge Distillation Framework for Recommender System
- MVIN: Learning Multiview Items for Recommendation
- Content-aware Neural Hashing for Cold-start Recommendation
- Signed Distance-based Deep Memory Recommender
- Learning the Structure of Auto-Encoding Recommenders
- On Sampling Collaborative Filtering Datasets
- DBRec: Dual-Bridging Recommendation via Discovering Latent Groups
- E-commerce Recommendation with Weighted Expected Utility
Cited by in corpus (6)
- Time to Split: Exploring Data Splitting Strategies for Offline Evaluation of Sequential Recommenders
- LightKG: Efficient Knowledge-Aware Recommendations with Simplified GNN Architecture
- A Thorough Performance Benchmarking on Lightweight Embedding-based Recommender Systems
- Beyond Collaborative Filtering: A Relook at Task Formulation in Recommender Systems
- Informfully Recommenders -- Reproducibility Framework for Diversity-aware Intra-session Recommendations
- APS Explorer: Navigating Algorithm Performance Spaces for Informed Dataset Selection