Graph-Based Global Reasoning Networks
arXiv:1811.12814
Abstract
Globally modeling and reasoning over relations between regions can be beneficial for many computer vision tasks on both images and videos. Convolutional Neural Networks (CNNs) excel at modeling local relations by convolution operations, but they are typically inefficient at capturing global relations between distant regions and require stacking multiple convolution layers. In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be efficiently computed. After reasoning, relation-aware features are distributed back to the original coordinate space for down-stream tasks. We further present a highly efficient instantiation of the proposed approach and introduce the Global Reasoning unit (GloRe unit) that implements the coordinate-interaction space mapping by weighted global pooling and weighted broadcasting, and the relation reasoning via graph convolution on a small graph in interaction space. The proposed GloRe unit is lightweight, end-to-end trainable and can be easily plugged into existing CNNs for a wide range of tasks. Extensive experiments show our GloRe unit can consistently boost the performance of state-of-the-art backbone architectures, including ResNet, ResNeXt, SE-Net and DPN, for both 2D and 3D CNNs, on image classification, semantic segmentation and video action recognition task.
References in corpus (18)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Deep Residual Learning for Image Recognition
- Neural Architecture Search with Reinforcement Learning
- The Kinetics Human Action Video Dataset
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- ParseNet: Looking Wider to See Better
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- A simple neural network module for relational reasoning
- Deformable Convolutional Networks
- Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning
- Aggregated Residual Transformations for Deep Neural Networks
- Pyramid Scene Parsing Network
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- Non-local Neural Networks
- Dual Path Networks
- Convolutional Random Walk Networks for Semantic Image Segmentation
Cited by in corpus (8)
- Attention Mechanisms in Computer Vision: A Survey
- Hierarchical Multi-Scale Attention for Semantic Segmentation
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution
- Dual Graph Convolutional Network for Semantic Segmentation
- Spatial Pyramid Based Graph Reasoning for Semantic Segmentation
- Deep Reasoning with Multi-Scale Context for Salient Object Detection
- Adaptively Connected Neural Networks
- Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition