Random Offset Block Embedding Array (ROBE) for CriteoTB Benchmark MLPerf DLRM Model : 1000 Compression and 3.1 Faster Inference
arXiv:2108.02191
Abstract
Deep learning for recommendation data is one of the most pervasive and challenging AI workload in recent times. State-of-the-art recommendation models are one of the largest models matching the likes of GPT-3 and Switch Transformer. Challenges in deep learning recommendation models (DLRM) stem from learning dense embeddings for each of the categorical tokens. These embedding tables in industrial scale models can be as large as hundreds of terabytes. Such large models lead to a plethora of engineering challenges, not to mention prohibitive communication overheads, and slower training and inference times. Of these, slower inference time directly impacts user experience. Model compression for DLRM is gaining traction and the community has recently shown impressive compression results. In this paper, we present Random Offset Block Embedding Array (ROBE) as a low memory alternative to embedding tables which provide orders of magnitude reduction in memory usage while maintaining accuracy and boosting execution speed. ROBE is a simple fundamental approach in improving both cache performance and the variance of randomized hashing, which could be of independent interest in itself. We demonstrate that we can successfully train DLRM models with same accuracy while using less memory. A compressed model directly results in faster inference without any engineering effort. In particular, we show that we can train DLRM model using ROBE array of size 100MB on a single GPU to achieve AUC of 0.8025 or higher as required by official MLPerf CriteoTB benchmark DLRM model of 100GB while achieving about (209\%) improvement in inference throughput.
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Efficient Estimation of Word Representations in Vector Space
- xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems
- Compressing Neural Networks with the Hashing Trick
- DeepFM: A Factorization-Machine based Neural Network for CTR Prediction
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems
- Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation Systems
- AutoEmb: Automated Embedding Dimensionality Search in Streaming Recommendations
- Differentiable Neural Input Search for Recommender Systems
- Learnable Embedding Sizes for Recommender Systems
- Semantically Constrained Memory Allocation (SCMA) for Embedding in Efficient Recommendation Systems