Language-Conditioned Graph Networks for Relational Reasoning
arXiv:1905.04405
Abstract
Solving grounded language tasks often requires reasoning about relationships between objects in the context of a given task. For example, to answer the question "What color is the mug on the plate?" we must check the color of the specific mug that satisfies the "on" relationship with respect to the plate. Recent work has proposed various methods capable of complex relational reasoning. However, most of their power is in the inference structure, while the scene is represented with simple local appearance features. In this paper, we take an alternate approach and build contextualized representations for objects in a visual scene to support relational reasoning. We propose a general framework of Language-Conditioned Graph Networks (LCGN), where each node represents an object, and is described by a context-aware representation from related objects through iterative message passing conditioned on the textual input. E.g., conditioning on the "on" relationship to the plate, the object "mug" gathers messages from the object "plate" to update its representation to "mug on the plate", which can be easily consumed by a simple classifier for answer prediction. We experimentally show that our LCGN approach effectively supports relational reasoning and improves performance across several tasks and datasets. Our code is available at http://ronghanghu.com/lcgn.
References in corpus (6)
- Adam: A Method for Stochastic Optimization
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Neural Message Passing for Quantum Chemistry
- Graph Neural Networks: A Review of Methods and Applications
- Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Cited by in corpus (9)
- Improving Multi-hop Knowledge Base Question Answering by Learning Intermediate Supervision Signals
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Unshuffling Data for Improved Generalization
- Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning
- Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
- Extracting Multiple-Relations in One-Pass with Pre-Trained Transformers
- Improving Question Answering over Incomplete KBs with Knowledge-Aware Reader
- Learning Cross-modal Context Graph for Visual Grounding
- Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks