Understanding the Effective Receptive Field in Deep Convolutional Neural Networks
arXiv:1701.04128
Abstract
We study characteristics of receptive fields of units in deep convolutional networks. The receptive field size is a crucial issue in many visual tasks, as the output must respond to large enough areas in the image to capture information about large objects. We introduce the notion of an effective receptive field, and show that it both has a Gaussian distribution and only occupies a fraction of the full theoretical receptive field. We analyze the effective receptive field in several architecture designs, and the effect of nonlinear activations, dropout, sub-sampling and skip connections on it. This leads to suggestions for ways to address its tendency to be too small.
Cited by in corpus (28)
- On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- Video Frame Interpolation via Adaptive Separable Convolution
- Neural Network-Based Automatic Liver Tumor Segmentation With Random Forest-Based Candidate Filtering
- LatentGNN: Learning Efficient Non-local Relations for Visual Recognition
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- SFD: Single Shot Scale-invariant Face Detector
- Computation Reallocation for Object Detection
- Structured Attentions for Visual Question Answering
- Blurring the Line Between Structure and Learning to Optimize and Adapt Receptive Fields
- Hyperspectral Image Classification With Context-Aware Dynamic Graph Convolutional Network
- FeatherNets: Convolutional Neural Networks as Light as Feather for Face Anti-spoofing
- Domain-invariant Stereo Matching Networks
- Irregular Convolutional Neural Networks
- Multi-Resolution Fully Convolutional Neural Networks for Monaural Audio Source Separation
- Learning from Videos with Deep Convolutional LSTM Networks
- AFO-TAD: Anchor-free One-Stage Detector for Temporal Action Detection
- Recurrent Attention Model with Log-Polar Mapping is Robust against Adversarial Attacks
- Subjective and Objective De-raining Quality Assessment Towards Authentic Rain Image
- 3D Neighborhood Convolution: Learning Depth-Aware Features for RGB-D and RGB Semantic Segmentation
- Detecting Reflections by Combining Semantic and Instance Segmentation
- A Spectral Nonlocal Block for Neural Networks
- Adaptive Context Network for Scene Parsing
- Encoder-Decoder based CNN and Fully Connected CRFs for Remote Sensed Image Segmentation
- Scaling up deep neural networks: a capacity allocation perspective
- Extra Proximal-Gradient Inspired Non-local Network
- Multi-Scale Convolutions for Learning Context Aware Feature Representations
- SAFE: Scale Aware Feature Encoder for Scene Text Recognition