Detecting Visual Relationships with Deep Relational Networks
arXiv:1704.03114
Abstract
Relationships among objects play a crucial role in image understanding. Despite the great success of deep learning techniques in recognizing individual objects, reasoning about the relationships among objects remains a challenging task. Previous methods often treat this as a classification problem, considering each type of relationship (e.g. "ride") or each distinct visual phrase (e.g. "person-ride-horse") as a category. Such approaches are faced with significant difficulties caused by the high diversity of visual appearance for each kind of relationships or the large number of distinct visual phrases. We propose an integrated framework to tackle this problem. At the heart of this framework is the Deep Relational Network, a novel formulation designed specifically for exploiting the statistical dependencies between objects and their relationships. On two large datasets, the proposed method achieves substantial improvement over state-of-the-art.
To be appeared in CVPR 2017 as an oral paper
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Fully Connected Deep Structured Networks
- Places: An Image Database for Deep Scene Understanding
- Visual Relationship Detection with Language Priors
- Visual Question Answering: A Survey of Methods and Datasets
- SPICE: Semantic Propositional Image Caption Evaluation
Cited by in corpus (32)
- The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
- Pixels to Graphs by Associative Embedding
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Towards Diverse and Natural Image Descriptions via a Conditional GAN
- Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
- Exploring Visual Relationship for Image Captioning
- Learning to Compose Dynamic Tree Structures for Visual Contexts
- Scene Graph Generation from Objects, Phrases and Region Captions
- Learning to Detect Human-Object Interactions
- Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering
- Natural Language Guided Visual Relationship Detection
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Tackling the Challenges in Scene Graph Generation with Local-to-Global Interactions
- Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition
- Grounding Referring Expressions in Images by Variational Context
- Object Relation Detection Based on One-shot Learning
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Graph R-CNN for Scene Graph Generation
- An Interpretable Model for Scene Graph Generation
- Relational Action Forecasting
- Scene Graph Parsing as Dependency Parsing
- Move Forward and Tell: A Progressive Generator of Video Descriptions
- LinkNet: Relational Embedding for Scene Graph
- 2nd Place Solution to the GQA Challenge 2019
- Grounded Objects and Interactions for Video Captioning
- Learning from the Scene and Borrowing from the Rich: Tackling the Long Tail in Scene Graph Generation
- ORD: Object Relationship Discovery for Visual Dialogue Generation
- Understand, Compose and Respond - Answering Visual Questions by a Composition of Abstract Procedures
- Learning Actor Relation Graphs for Group Activity Recognition
- Tell-the-difference: Fine-grained Visual Descriptor via a Discriminating Referee
- Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features
- Adaptive and Iteratively Improving Recurrent Lateral Connections