Multi-Modal Self-Supervised Learning for Recommendation
arXiv:2302.10632 · doi:10.1145/3543507.3583206
Abstract
The online emergence of multi-modal sharing platforms (eg, TikTok, Youtube) is powering personalized recommender systems to incorporate various modalities (eg, visual, textual and acoustic) into the latent user representations. While existing works on multi-modal recommendation exploit multimedia content features in enhancing item embeddings, their model representation capability is limited by heavy label reliance and weak robustness on sparse user behavior data. Inspired by the recent progress of self-supervised learning in alleviating label scarcity issue, we explore deriving self-supervision signals with effectively learning of modality-aware user preference and cross-modal dependencies. To this end, we propose a new Multi-Modal Self-Supervised Learning (MMSSL) method which tackles two key challenges. Specifically, to characterize the inter-dependency between the user-item collaborative view and item multi-modal semantic view, we design a modality-aware interactive structure learning paradigm via adversarial perturbations for data augmentation. In addition, to capture the effects that user's modality-aware interaction pattern would interweave with each other, a cross-modal contrastive learning approach is introduced to jointly preserve the inter-modal semantic commonality and user preference diversity. Experiments on real-world datasets verify the superiority of our method in offering great potential for multimedia recommendation over various state-of-the-art baselines. The implementation is released at: https://github.com/HKUDS/MMSSL.
This paper has been published as a full paper at WWW 2023
References in corpus (9)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Improving neural networks by preventing co-adaptation of feature detectors
- KGAT: Knowledge Graph Attention Network for Recommendation
- S^3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization
- Improving Graph Collaborative Filtering with Neighborhood-enriched Contrastive Learning
- Global Context Enhanced Graph Neural Networks for Session-based Recommendation
- Hypergraph Contrastive Collaborative Filtering
- Heterogeneous Graph Contrastive Learning for Recommendation
- Contrastive Meta Learning with Behavior Multiplicity for Recommendation
Cited by in corpus (13)
- Multi-View Graph Convolutional Network for Multimedia Recommendation
- Knowledge Graph Self-Supervised Rationalization for Recommendation
- Graph Transformer for Recommendation
- Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
- Exploring Adapter-based Transfer Learning for Recommender Systems: Empirical Studies and Practical Insights
- Formalizing Multimedia Recommendation through Multimodal Deep Learning
- IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT
- A Comprehensive Survey on Self-Supervised Learning for Recommendation
- Pareto Invariant Representation Learning for Multimedia Recommendation
- ABXI: Invariant Interest Adaptation for Task-Guided Cross-Domain Sequential Recommendation
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems
- I-MRec: Invariant Learning with Information Bottleneck for Incomplete Modality Recommendation
- Transferable and Forecastable User Targeting Foundation Model