Relation Networks for Object Detection
arXiv:1711.11575
Abstract
Although it is well believed for years that modeling relations between objects would help object recognition, there has not been evidence that the idea is working in the deep learning era. All state-of-the-art object detection systems still rely on recognizing object instances individually, without exploiting their relations during learning. This work proposes an object relation module. It processes a set of objects simultaneously through interaction between their appearance feature and geometry, thus allowing modeling of their relations. It is lightweight and in-place. It does not require additional supervision and is easy to embed in existing networks. It is shown effective on improving object recognition and duplicate removal steps in the modern object detection pipeline. It verifies the efficacy of modeling object relations in CNN based detection. It gives rise to the first fully end-to-end object detector.
References in corpus (6)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Interaction Networks for Learning about Objects, Relations and Physics
- A simple neural network module for relational reasoning
- Discovering objects and their relations from entangled scene representations
- Visual Interaction Networks
- Massive Exploration of Neural Machine Translation Architectures
Cited by in corpus (20)
- Dual Attention Network for Scene Segmentation
- Deformable ConvNets v2: More Deformable, Better Results
- Relational Deep Reinforcement Learning
- Deep Closest Point: Learning Representations for Point Cloud Registration
- Acquisition of Localization Confidence for Accurate Object Detection
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Relational recurrent neural networks
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- Multi-Task Deep Networks for Depth-Based 6D Object Pose and Joint Registration in Crowd Scenarios
- Solution for Large-Scale Hierarchical Object Detection Datasets with Incomplete Annotation and Data Imbalance
- Relation Networks for Optic Disc and Fovea Localization in Retinal Images
- Disentangled Non-Local Neural Networks
- MultiResolution Attention Extractor for Small Object Detection
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
- Deep Regionlets for Object Detection
- Relational Action Forecasting
- ASAP-NMS: Accelerating Non-Maximum Suppression Using Spatially Aware Priors
- Learning Actor Relation Graphs for Group Activity Recognition
- G-RCN: Optimizing the Gap between Classification and Localization Tasks for Object Detection
- Object Detection on Single Monocular Images through Canonical Correlation Analysis