Embedding Compression in Recommender Systems: A Survey
arXiv:2408.02304 · doi:10.1145/3637841
Abstract
To alleviate the problem of information explosion, recommender systems are widely deployed to provide personalized information filtering services. Usually, embedding tables are employed in recommender systems to transform high-dimensional sparse one-hot vectors into dense real-valued embeddings. However, the embedding tables are huge and account for most of the parameters in industrial-scale recommender systems. In order to reduce memory costs and improve efficiency, various approaches are proposed to compress the embedding tables. In this survey, we provide a comprehensive review of embedding compression approaches in recommender systems. We first introduce deep learning recommendation models and the basic concept of embedding compression in recommender systems. Subsequently, we systematically organize existing approaches into three categories, namely low-precision, mixed-dimension, and weight-sharing, respectively. Lastly, we summarize the survey with some general suggestions and provide future prospects for this field.
Accepted by ACM Computing Surveys
References in corpus (17)
- AutoML: A Survey of the State-of-the-Art
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems
- A Survey on Accuracy-oriented Neural Recommendation: From Collaborative Filtering to Information-rich Recommendation
- Feature Generation by Convolutional Neural Network for Click-Through Rate Prediction
- Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for Recommendation
- OptEmbed: Learning Optimal Embedding Table for Click-through Rate Prediction
- Detecting Arbitrary Order Beneficial Feature Interactions for Recommender Systems
- Differentiable Neural Input Search for Recommender Systems
- Adaptive Low-Precision Training for Embeddings in Click-Through Rate Prediction
- Memory-efficient Embedding for Recommendations
- Linear-Time Self Attention with Codeword Histogram for Efficient Recommendation
- Post-Training 4-bit Quantization on Embedding Tables
- Mixed-Precision Embedding Using a Cache
- Training with Multi-Layer Embeddings for Model Reduction
- Semantically Constrained Memory Allocation (SCMA) for Embedding in Efficient Recommendation Systems
- Towards Low-loss 1-bit Quantization of User-item Representations for Top-K Recommendation