Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering
arXiv:1708.03619 · doi:10.1109/TNNLS.2018.2817340
Abstract
Visual question answering (VQA) is challenging because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multi-modal feature fusion that is able to capture the complex interactions between multi-modal features; 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a `co-attention' mechanism is developed by using a deep neural network architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multi-modal feature fusion, a generalized Multi-modal Factorized High-order pooling approach (MFH) is developed to achieve more effective fusion of multi-modal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the KL (Kullback-Leibler) divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A deep neural network architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA datasets and win the runner-up in VQA Challenge 2017.
13 pages, 9 figures. arXiv admin note: substantial text overlap with arXiv:1708.01471
References in corpus (7)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- A Structured Self-attentive Sentence Embedding
- Skip-Thought Vectors
- A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
- Hadamard Product for Low-rank Bilinear Pooling
- Dual Attention Networks for Multimodal Reasoning and Matching
- Learning Convolutional Text Representations for Visual Question Answering
Cited by in corpus (66)
- Multimodal Intelligence: Representation Learning, Information Fusion, and Applications
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers
- RUBi: Reducing Unimodal Biases in Visual Question Answering
- Learning to Count Objects in Natural Images for Visual Question Answering
- Medical Visual Question Answering: A Survey
- Pythia v0.1: the Winning Entry to the VQA Challenge 2018
- Biomedical Question Answering: A Survey of Approaches and Challenges
- Deep Modular Co-Attention Networks for Visual Question Answering
- AU R-CNN: Encoding Expert Prior Knowledge into R-CNN for Action Unit Detection
- DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- From Easy to Hard: Learning Language-guided Curriculum for Visual Question Answering on Remote Sensing Data
- Optimizing Deep Neural Network Architecture: A Tabu Search Based Approach
- Towards Lightweight Transformer via Group-wise Transformation for Vision-and-Language Tasks
- A scoping review on multimodal deep learning in biomedical images and texts
- Relation-Aware Graph Attention Network for Visual Question Answering
- Mixed High-Order Attention Network for Person Re-Identification
- An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
- Multimodal Unified Attention Networks for Vision-and-Language Interactions
- Multi-modality Latent Interaction Network for Visual Question Answering
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- A Dual-Attention Learning Network with Word and Sentence Embedding for Medical Visual Question Answering
- Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection
- BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
- Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering
- RTIC: Residual Learning for Text and Image Composition using Graph Convolutional Network
- MUREL: Multimodal Relational Reasoning for Visual Question Answering
- In Defense of Grid Features for Visual Question Answering
- Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism
- Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
- How to find a good image-text embedding for remote sensing visual question answering?
- Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text
- Joint Image Captioning and Question Answering
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- VrR-VG: Refocusing Visually-Relevant Relationships
- Deep Multimodal Neural Architecture Search
- Explainable High-order Visual Question Reasoning: A New Benchmark and Knowledge-routed Network
- MuVAM: A Multi-View Attention-based Model for Medical Visual Question Answering
- Learning Rich Image Region Representation for Visual Question Answering
- Self-Segregating and Coordinated-Segregating Transformer for Focused Deep Multi-Modular Network for Visual Question Answering
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual Grounding
- Dual Recurrent Attention Units for Visual Question Answering
- Stroke Constrained Attention Network for Online Handwritten Mathematical Expression Recognition
- Question-Agnostic Attention for Visual Question Answering
- Bilinear Graph Networks for Visual Question Answering
- Knowledge-Routed Visual Question Reasoning: Challenges for Deep Representation Embedding
- Regularizing Attention Networks for Anomaly Detection in Visual Question Answering
- Multimodal Learning for Hateful Memes Detection
- ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration
- New Ideas and Trends in Deep Multimodal Content Understanding: A Review
- An Entropy Clustering Approach for Assessing Visual Question Difficulty
- SurgicalGPT: End-to-End Language-Vision GPT for Visual Question Answering in Surgery
- DNN-based cross-lingual voice conversion using Bottleneck Features
- 6G: Connecting Everything by 1000 Times Price Reduction
- Frontal Low-rank Random Tensors for Fine-grained Action Segmentation
- Geometry-Entangled Visual Semantic Transformer for Image Captioning
- Trying Bilinear Pooling in Video-QA
- The Resale Price Prediction of Secondhand Jewelry Items Using a Multi-modal Deep Model with Iterative Co-Attention
- Visual Question Answering based on Local-Scene-Aware Referring Expression Generation
- Interpretable Visual Question Answering by Visual Grounding from Attention Supervision Mining
- Adaptively Denoising Proposal Collection for Weakly Supervised Object Localization
- Modulated Self-attention Convolutional Network for VQA
- Learning to Represent and Predict Sets with Deep Neural Networks
- Discovering the Unknown Knowns: Turning Implicit Knowledge in the Dataset into Explicit Training Examples for Visual Question Answering
- Phrase Grounding by Soft-Label Chain Conditional Random Field
- After All, Only The Last Neuron Matters: Comparing Multi-modal Fusion Functions for Scene Graph Generation