Select-Additive Learning: Improving Generalization in Multimodal Sentiment Analysis
arXiv:1609.05244
Abstract
Multimodal sentiment analysis is drawing an increasing amount of attention these days. It enables mining of opinions in video reviews which are now available aplenty on online platforms. However, multimodal sentiment analysis has only a few high-quality data sets annotated for training machine learning algorithms. These limited resources restrict the generalizability of models, where, for example, the unique characteristics of a few speakers (e.g., wearing glasses) may become a confounding factor for the sentiment classification task. In this paper, we propose a Select-Additive Learning (SAL) procedure that improves the generalizability of trained neural networks for multimodal sentiment analysis. In our experiments, we show that our SAL approach improves prediction accuracy significantly in all three modalities (verbal, acoustic, visual), as well as in their fusion. Our results show that SAL, even when trained on one dataset, achieves good generalization across two new test datasets.
Supplementary files at: http://www.cs.cmu.edu/~haohanw/document/sal_supp.pdf
References in corpus (1)
Cited by in corpus (7)
- On the Origin of Deep Learning
- Efficient Low-rank Multimodal Fusion with Modality-Specific Factors
- Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors
- Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities
- Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment
- What If We Simply Swap the Two Text Fragments? A Straightforward yet Effective Way to Test the Robustness of Methods to Confounding Signals in Nature Language Inference Tasks
- Multimodal Sentiment Analysis with Word-Level Fusion and Reinforcement Learning