Learning Conditioned Graph Structures for Interpretable Visual Question Answering
arXiv:1806.07243
Abstract
Visual Question answering is a challenging problem requiring a combination of concepts from Computer Vision and Natural Language Processing. Most existing approaches use a two streams strategy, computing image and question features that are consequently merged using a variety of techniques. Nonetheless, very few rely on higher level image representations, which can capture semantic and spatial relationships. In this paper, we propose a novel graph-based approach for Visual Question Answering. Our method combines a graph learner module, which learns a question specific graph representation of the input image, with the recent concept of graph convolutions, aiming to learn image representations that capture question specific interactions. We test our approach on the VQA v2 dataset using a simple baseline architecture enhanced by the proposed graph learner module. We obtain promising results with 66.18% accuracy and demonstrate the interpretability of the proposed method. Code can be found at github.com/aimbrain/vqa-project.
NIPS 2018 (13 pages, 7 figures)
Cited by in corpus (36)
- Graph Neural Networks: A Review of Methods and Applications
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Large-Scale Adversarial Training for Vision-and-Language Representation Learning
- Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node Embeddings
- Differentiable Graph Module (DGM) for Graph Convolutional Networks
- Cross-modal Knowledge Reasoning for Knowledge-based Visual Question Answering
- Multi-modal Deep Analysis for Multimedia
- Relation-Aware Graph Attention Network for Visual Question Answering
- Deep Iterative and Adaptive Learning for Graph Neural Networks
- Object Relational Graph with Teacher-Recommended Learning for Video Captioning
- Multi-modality Latent Interaction Network for Visual Question Answering
- GraphFlow: Exploiting Conversation Flow with Graph Neural Networks for Conversational Machine Comprehension
- MUREL: Multimodal Relational Reasoning for Visual Question Answering
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- Language and Visual Entity Relationship Graph for Agent Navigation
- Intrinsic Relationship Reasoning for Small Object Detection
- Counterfactual Critic Multi-Agent Training for Scene Graph Generation
- Learning by Abstraction: The Neural State Machine
- Linguistically-aware Attention for Reducing the Semantic-Gap in Vision-Language Tasks
- Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text
- VrR-VG: Refocusing Visually-Relevant Relationships
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Cascaded Human-Object Interaction Recognition
- Graph Structured Network for Image-Text Matching
- CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation
- Answer Them All! Toward Universal Visual Question Answering Models
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation
- Location-aware Graph Convolutional Networks for Video Question Answering
- GraphSearchNet: Enhancing GNNs via Capturing Global Dependencies for Semantic Code Search
- Learning Contextualized Knowledge Structures for Commonsense Reasoning
- Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension
- A Peek Into the Reasoning of Neural Networks: Interpreting with Structural Visual Concepts
- Exploiting Relationship for Complex-scene Image Generation
- Non-monotonic Logical Reasoning Guiding Deep Learning for Explainable Visual Question Answering