Pareto Invariant Representation Learning for Multimedia Recommendation
arXiv:2308.04706 · doi:10.1145/3581783.3612591
Abstract
Multimedia recommendation involves personalized ranking tasks, where multimedia content is usually represented using a generic encoder. However, these generic representations introduce spurious correlations that fail to reveal users' true preferences. Existing works attempt to alleviate this problem by learning invariant representations, but overlook the balance between independent and identically distributed (IID) and out-of-distribution (OOD) generalization. In this paper, we propose a framework called Pareto Invariant Representation Learning (PaInvRL) to mitigate the impact of spurious correlations from an IID-OOD multi-objective optimization perspective, by learning invariant representations (intrinsic factors that attract user attention) and variant representations (other factors) simultaneously. Specifically, PaInvRL includes three iteratively executed modules: (i) heterogeneous identification module, which identifies the heterogeneous environments to reflect distributional shifts for user-item interactions; (ii) invariant mask generation module, which learns invariant masks based on the Pareto-optimal solutions that minimize the adaptive weighted Invariant Risk Minimization (IRM) and Empirical Risk (ERM) losses; (iii) convert module, which generates both variant representations and item-invariant representations for training a multi-modal recommendation model that mitigates spurious correlations and balances the generalization performance within and cross the environmental distributions. We compare the proposed PaInvRL with state-of-the-art recommendation models on three public multimedia recommendation datasets (Movielens, Tiktok, and Kwai), and the experimental results validate the effectiveness of PaInvRL for both within- and cross-environmental learning.
ACM MM 2023 full paper
References in corpus (29)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Graph Collaborative Filtering
- ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
- BART: Bayesian additive regression trees
- Session-based Recommendation with Graph Neural Networks
- Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
- LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation
- Mining Latent Structures for Multimedia Recommendation
- Causal Embeddings for Recommendation
- Multi-Modal Self-Supervised Learning for Recommendation
- Collaborative Filtering and the Missing at Random Assumption
- Aesthetic-based Clothing Recommendation
- AutoDebias: Learning to Debias for Recommendation
- Bias and Debias in Recommender System: A Survey and Future Directions
- User Diverse Preference Modeling by Multimodal Attentive Metric Learning
- ESCM: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation
- Enhanced Doubly Robust Learning for Debiasing Post-click Conversion Rate Estimation
- Perfect Match: A Simple Method for Learning Representations For Counterfactual Inference With Neural Networks
- Balancing Unobserved Confounding with a Few Unbiased Ratings in Debiased Recommendations
- UltraGCN: Ultra Simplification of Graph Convolutional Networks for Recommendation
- DESCN: Deep Entire Space Cross Networks for Individual Treatment Effect Estimation
- ZIN: When and How to Learn Invariance Without Environment Partition?
- Towards Domain Generalization in Object Detection
- TDR-CL: Targeted Doubly Robust Collaborative Learning for Debiased Recommendations
- StableDR: Stabilized Doubly Robust Learning for Recommendation on Data Missing Not at Random
- Regulatory Instruments for Fair Personalized Pricing
- On the Opportunity of Causal Learning in Recommendation Systems: Foundation, Estimation, Prediction and Challenges
- Kernelized Heterogeneous Risk Minimization
- Learning Representations that Support Robust Transfer of Predictors