Turning Dross Into Gold Loss: is BERT4Rec really better than SASRec?
arXiv:2309.07602 · doi:10.1145/3604915.3610644
Abstract
Recently sequential recommendations and next-item prediction task has become increasingly popular in the field of recommender systems. Currently, two state-of-the-art baselines are Transformer-based models SASRec and BERT4Rec. Over the past few years, there have been quite a few publications comparing these two algorithms and proposing new state-of-the-art models. In most of the publications, BERT4Rec achieves better performance than SASRec. But BERT4Rec uses cross-entropy over softmax for all items, while SASRec uses negative sampling and calculates binary cross-entropy loss for one positive and one negative item. In our work, we show that if both models are trained with the same loss, which is used by BERT4Rec, then SASRec will significantly outperform BERT4Rec both in terms of quality and training speed. In addition, we show that SASRec could be effectively trained with negative sampling and still outperform BERT4Rec, but the number of negative examples should be much larger than one.
References in corpus (6)
- Contrastive Learning for Representation Degeneration Problem in Sequential Recommendation
- Translation-based Recommendation
- Lightweight Self-Attentive Sequential Recommendation
- Contrastive Self-supervised Sequential Recommendation with Robust Augmentation
- Denoising Self-attentive Sequential Recommendation
- A Case Study on Sampling Strategies for Evaluating Neural Sequential Item Recommendation Models
Cited by in corpus (17)
- Does It Look Sequential? An Analysis of Datasets for Evaluation of Sequential Recommendations
- A Reproducible Analysis of Sequential Recommender Systems
- Time to Split: Exploring Data Splitting Strategies for Offline Evaluation of Sequential Recommenders
- Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item Catalogs
- Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach
- RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
- RePlay: a Recommendation Framework for Experimentation and Production Use
- Calibration-Disentangled Learning and Relevance-Prioritized Reranking for Calibrated Sequential Recommendation
- Efficient Recommendation with Millions of Items by Dynamic Pruning of Sub-Item Embeddings
- Autoregressive Generation Strategies for Top-K Sequential Recommendations
- Recommendation Is a Dish Better Served Warm
- Efficient Inference of Sub-Item Id-based Sequential Recommendation Models with Millions of Items
- eSASRec: Enhancing Transformer-based Recommendations in a Modular Fashion
- Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
- Correcting the LogQ Correction: Revisiting Sampled Softmax for Large-Scale Retrieval
- Multi-Item-Query Attention for Stable Sequential Recommendation
- Exploiting Preferences in Loss Functions for Sequential Recommendation via Weak Transitivity