Multimodal Machine Learning: A Survey and Taxonomy
arXiv:1705.09406
Abstract
Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is characterized as multimodal when it includes multiple such modalities. In order for Artificial Intelligence to make progress in understanding the world around us, it needs to be able to interpret such multimodal signals together. Multimodal machine learning aims to build models that can process and relate information from multiple modalities. It is a vibrant multi-disciplinary field of increasing importance and with extraordinary potential. Instead of focusing on specific multimodal applications, this paper surveys the recent advances in multimodal machine learning itself and presents them in a common taxonomy. We go beyond the typical early and late fusion categorization and identify broader challenges that are faced by multimodal machine learning, namely: representation, translation, alignment, fusion, and co-learning. This new taxonomy will enable researchers to better understand the state of the field and identify directions for future research.
References in corpus (16)
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- WaveNet: A Generative Model for Raw Audio
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Dynamic Memory Networks for Visual and Textual Question Answering
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- Show and Tell: A Neural Image Caption Generator
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion
- Language Models for Image Captioning: The Quirks and What Works
- Multi-View Learning in the Presence of View Disagreement
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Deep Multimodal Learning for Audio-Visual Speech Recognition
Cited by in corpus (15)
- Recent Trends in Deep Learning Based Natural Language Processing
- Learning Factorized Multimodal Representations
- Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities
- Multimodal deep learning for short-term stock volatility prediction
- Responsible and Representative Multimodal Data Acquisition and Analysis: On Auditability, Benchmarking, Confidence, Data-Reliance & Explainability
- Improved Robust ASR for Social Robots in Public Spaces
- Variational Selective Autoencoder: Learning from Partially-Observed Heterogeneous Data
- Multimodal Sentiment Analysis with Word-Level Fusion and Reinforcement Learning
- An Efficient Approach for Geo-Multimedia Cross-Modal Retrieval
- Multi-Source Neural Variational Inference
- Quantum-inspired Multimodal Fusion for Video Sentiment Analysis
- Social Fraud Detection Review: Methods, Challenges and Analysis
- 3D-MOV: Audio-Visual LSTM Autoencoder for 3D Reconstruction of Multiple Objects from Video
- Contextual Fusion For Adversarial Robustness
- Interpretable, similarity-driven multi-view embeddings from high-dimensional biomedical data