AdaMCT: Adaptive Mixture of CNN-Transformer for Sequential Recommendation
arXiv:2205.08776 · doi:10.1145/3583780.3614773
Abstract
Sequential recommendation (SR) aims to model users dynamic preferences from a series of interactions. A pivotal challenge in user modeling for SR lies in the inherent variability of user preferences. An effective SR model is expected to capture both the long-term and short-term preferences exhibited by users, wherein the former can offer a comprehensive understanding of stable interests that impact the latter. To more effectively capture such information, we incorporate locality inductive bias into the Transformer by amalgamating its global attention mechanism with a local convolutional filter, and adaptively ascertain the mixing importance on a personalized basis through layer-aware adaptive mixture units, termed as AdaMCT. Moreover, as users may repeatedly browse potential purchases, it is expected to consider multiple relevant items concurrently in long-/short-term preferences modeling. Given that softmax-based attention may promote unimodal activation, we propose the Squeeze-Excitation Attention (with sigmoid activation) into SR models to capture multiple pertinent items (keys) simultaneously. Extensive experiments on three widely employed benchmarks substantiate the effectiveness and efficiency of our proposed approach. Source code is available at https://github.com/juyongjiang/AdaMCT.
Accepted by CIKM 2023
References in corpus (6)
- Translation-based Recommendation
- STAN: Spatio-Temporal Attention Network for Next Location Recommendation
- TAGNN: Target Attentive Graph Neural Networks for Session-based Recommendation
- Efficiently Leveraging Multi-level User Intent for Session-based Recommendation via Atten-Mixer Network
- Evolutionary Preference Learning via Graph Nested GRU ODE for Session-based Recommendation
- An Adaptive Graph Pre-training Framework for Localized Collaborative Filtering