Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions
arXiv:2107.13782 · doi:10.1016/j.inffus.2021.12.003
Abstract
Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves multiple aspects: representation, translation, alignment, fusion, and co-learning. In the current state of multimodal machine learning, the assumptions are that all modalities are present, aligned, and noiseless during training and testing time. However, in real-world tasks, typically, it is observed that one or more modalities are missing, noisy, lacking annotated data, have unreliable labels, and are scarce in training or testing and or both. This challenge is addressed by a learning paradigm called multimodal co-learning. The modeling of a (resource-poor) modality is aided by exploiting knowledge from another (resource-rich) modality using transfer of knowledge between modalities, including their representations and predictive models. Co-learning being an emerging area, there are no dedicated reviews explicitly focusing on all challenges addressed by co-learning. To that end, in this work, we provide a comprehensive survey on the emerging area of multimodal co-learning that has not been explored in its entirety yet. We review implementations that overcome one or more co-learning challenges without explicitly considering them as co-learning challenges. We present the comprehensive taxonomy of multimodal co-learning based on the challenges addressed by co-learning and associated implementations. The various techniques employed to include the latest ones are reviewed along with some of the applications and datasets. Our final goal is to discuss challenges and perspectives along with the important ideas and directions for future work that we hope to be beneficial for the entire research community focusing on this exciting domain.
This is published in Information Fusion Journal and published copy can be downloaded from https://authors.elsevier.com/c/1eIVS5a7-Gls0Y for 50 days. https://www.sciencedirect.com/science/article/pii/S1566253521002530
References in corpus (29)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Conditional Generative Adversarial Nets
- Language Models are Few-Shot Learners
- Domain Generalization: A Survey
- Zero-Shot Learning Through Cross-Modal Transfer
- Multi-Task Learning with Deep Neural Networks: A Survey
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Billion-scale semi-supervised learning for image classification
- WebVision Database: Visual Learning and Understanding from Web Data
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
- Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
- A practical tutorial on autoencoders for nonlinear feature fusion: Taxonomy, models, software and guidelines
- An Overview of Deep Semi-Supervised Learning
- Heterogeneous Domain Generalization via Domain Mixup
- Gas Detection and Identification Using Multimodal Artificial Intelligence Based Sensor Fusion
- A Transformer-based joint-encoding for Emotion Recognition and Sentiment Analysis
- Self-training Improves Pre-training for Natural Language Understanding
- A Closer Look at the Robustness of Vision-and-Language Pre-trained Models
- Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges
- Enriched Music Representations with Multiple Cross-modal Contrastive Learning
- Unsupervised Audio-Visual Subspace Alignment for High-Stakes Deception Detection
- Contrastive Visual-Linguistic Pretraining
- SMIL: Multimodal Learning with Severely Missing Modality
- Question-Conditioned Counterfactual Image Generation for VQA
- MCQA: Multimodal Co-attention Based Network for Question Answering
- CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual Generation
- COBRA: Contrastive Bi-Modal Representation Algorithm
- Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention
- MMED: A Multi-domain and Multi-modality Event Dataset
Cited by in corpus (5)
- Visuo-Haptic Object Perception for Robots: An Overview
- Patchwork Learning: A Paradigm Towards Integrative Analysis across Diverse Biomedical Data Sources
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Interpreting wealth distribution via poverty map inference using multimodal data
- Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration