KuaiRec: A Fully-observed Dataset and Insights for Evaluating Recommender Systems
arXiv:2202.10842 · doi:10.1145/3511808.3557220
Abstract
The progress of recommender systems is hampered mainly by evaluation as it requires real-time interactions between humans and systems, which is too laborious and expensive. This issue is usually approached by utilizing the interaction history to conduct offline evaluation. However, existing datasets of user-item interactions are partially observed, leaving it unclear how and to what extent the missing interactions will influence the evaluation. To answer this question, we collect a fully-observed dataset from Kuaishou's online environment, where almost all 1,411 users have been exposed to all 3,327 items. To the best of our knowledge, this is the first real-world fully-observed data with millions of user-item interactions. With this unique dataset, we conduct a preliminary analysis of how the two factors - data density and exposure bias - affect the evaluation results of multi-round conversational recommendation. Our main discoveries are that the performance ranking of different methods varies with the two factors, and this effect can only be alleviated in certain cases by estimating missing interactions for user simulation. This demonstrates the necessity of the fully-observed dataset. We release the dataset and the pipeline implementation for evaluation at https://kuairec.com
CIKM '22 Full Paper
References in corpus (10)
- S^3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization
- Causal Intervention for Leveraging Popularity Bias in Recommendation
- Estimation-Action-Reflection: Towards Deep Interaction Between Conversational and Recommender Systems
- Collaborative Filtering and the Missing at Random Assumption
- Offline A/B testing for Recommender Systems
- Counterfactual Risk Minimization: Learning from Logged Bandit Feedback
- Evaluating Conversational Recommender Systems via User Simulation
- User-Centric Conversational Recommendation with Multi-Aspect User Modeling
- Reducing Offline Evaluation Bias in Recommendation Systems
- C2-CRS: Coarse-to-Fine Contrastive Learning for Conversational Recommender System
Cited by in corpus (16)
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive Recommendation
- SIGformer: Sign-aware Graph Transformer for Recommendation
- KuaiSAR: A Unified Search And Recommendation Dataset
- Going Beyond Popularity and Positivity Bias: Correcting for Multifactorial Bias in Recommender Systems
- From Variability to Stability: Advancing RecSys Benchmarking Practices
- Relevance meets Diversity: A User-Centric Framework for Knowledge Exploration through Recommendations
- ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender Systems
- Treatment Effect Estimation for User Interest Exploration on Recommender Systems
- Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems
- Beyond Static Calibration: The Impact of User Preference Dynamics on Calibrated Recommendation
- Multi-Granularity Distribution Modeling for Video Watch Time Prediction via Exponential-Gaussian Mixture Network
- Revisiting Clustering of Neural Bandits: Selective Reinitialization for Mitigating Loss of Plasticity
- On the Reliability of Sampling Strategies in Offline Recommender Evaluation
- PPVF: An Efficient Privacy-Preserving Online Video Fetching Framework with Correlated Differential Privacy
- LLM-Enhanced Reinforcement Learning for Long-Term User Satisfaction in Interactive Recommendation
- Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization