Beyond NDCG: behavioral testing of recommender systems with RecList
arXiv:2111.09963 · doi:10.1145/3487553.3524215
Abstract
As with most Machine Learning systems, recommender systems are typically evaluated through performance metrics computed over held-out data points. However, real-world behavior is undoubtedly nuanced: ad hoc error analysis and deployment-specific tests must be employed to ensure the desired quality in actual deployments. In this paper, we propose RecList, a behavioral-based testing methodology. RecList organizes recommender systems by use case and introduces a general plug-and-play procedure to scale up behavioral testing. We demonstrate its capabilities by analyzing known algorithms and black-box commercial systems, and we release RecList as an open source, extensible package for the community.
Paper accepted to the WebConf 2022
References in corpus (8)
- Meta-Prod2Vec - Product Embeddings Using Side-Information for Recommendation
- M2GRL: A Multi-task Multi-view Graph Representation Learning Framework for Web-scale Recommender Systems
- SIGIR 2021 E-Commerce Workshop Data Challenge
- Fantastic Embeddings and How to Align Them: Zero-Shot Inference in a Multi-Shop Scenario
- Transformers with multi-modal features and post-fusion context for e-commerce session-based recommendation
- You Do Not Need a Bigger Boat: Recommendations at Reasonable Scale in a (Mostly) Serverless and Open Stack
- "Are you sure?": Preliminary Insights from Scaling Product Comparisons to Multiple Shops
- DAG Card is the new Model Card