Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering
arXiv:1812.09681
Abstract
One of the key issues of Visual Question Answering (VQA) is to reason with semantic clues in the visual content under the guidance of the question, how to model relational semantics still remains as a great challenge. To fully capture visual semantics, we propose to reason over a structured visual representation - scene graph, with embedded objects and inter-object relationships. This shows great benefit over vanilla vector representations and implicit visual relationship learning. Based on existing visual relationship models, we propose a visual relationship encoder that projects visual relationships into a learned deep semantic space constrained by visual context and language priors. Upon the constructed graph, we propose a Scene Graph Convolutional Network (SceneGCN) to jointly reason the object properties and relational semantics for the correct answer. We demonstrate the model's effectiveness and interpretability on the challenging GQA dataset and the classical VQA 2.0 dataset, remarkably achieving state-of-the-art 54.56% accuracy on GQA compared to the existing best model.
14 pages, 9 figures
References in corpus (13)
- Semi-Supervised Classification with Graph Convolutional Networks
- Neural Message Passing for Quantum Chemistry
- Graph Neural Networks: A Review of Methods and Applications
- VQA: Visual Question Answering
- Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering
- Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
- Simple Baseline for Visual Question Answering
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Hadamard Product for Low-rank Bilinear Pooling
- Detecting Visual Relationships with Deep Relational Networks
- Identity-Aware Textual-Visual Matching with Latent Co-attention
- Finding Linear Structure in Large Datasets with Scalable Canonical Correlation Analysis
- A Simple Loss Function for Improving the Convergence and Accuracy of Visual Question Answering Models
Cited by in corpus (8)
- A Comprehensive Survey of Scene Graphs: Generation and Application
- Relation-Aware Graph Attention Network for Visual Question Answering
- Multi-modality Latent Interaction Network for Visual Question Answering
- An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
- Predicate correlation learning for scene graph generation
- Learning Visual Relation Priors for Image-Text Matching and Image Captioning with Neural Scene Graph Generators
- Character Matters: Video Story Understanding with Character-Aware Relations
- Scene Graph Reasoning for Visual Question Answering