Relation-Aware Graph Attention Network for Visual Question Answering
arXiv:1903.12314
Abstract
In order to answer semantically-complicated questions about an image, a Visual Question Answering (VQA) model needs to fully understand the visual scene in the image, especially the interactive dynamics between different objects. We propose a Relation-aware Graph Attention Network (ReGAT), which encodes each image into a graph and models multi-type inter-object relations via a graph attention mechanism, to learn question-adaptive relation representations. Two types of visual object relations are explored: (i) Explicit Relations that represent geometric positions and semantic interactions between objects; and (ii) Implicit Relations that capture the hidden dynamics between image regions. Experiments demonstrate that ReGAT outperforms prior state-of-the-art approaches on both VQA 2.0 and VQA-CP v2 datasets. We further show that ReGAT is compatible to existing VQA architectures, and can be used as a generic relation encoder to boost the model performance for VQA.
To appear in ICCV 2019
References in corpus (8)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering
- Learning to Count Objects in Natural Images for Visual Question Answering
- Hadamard Product for Low-rank Bilinear Pooling
- Pythia v0.1: the Winning Entry to the VQA Challenge 2018
- Learning Conditioned Graph Structures for Interpretable Visual Question Answering
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering
Cited by in corpus (9)
- Large-Scale Adversarial Training for Vision-and-Language Representation Learning
- Cross-modal Knowledge Reasoning for Knowledge-based Visual Question Answering
- An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
- Multi-step Reasoning via Recurrent Dual Attention for Visual Dialog
- Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
- Learning by Abstraction: The Neural State Machine
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
- New Ideas and Trends in Deep Multimodal Content Understanding: A Review
- Adventurer's Treasure Hunt: A Transparent System for Visually Grounded Compositional Visual Question Answering based on Scene Graphs