Understanding the Effective Receptive Field in Deep Convolutional Neural Networks
arXiv:1701.04128
Abstract
We study characteristics of receptive fields of units in deep convolutional networks. The receptive field size is a crucial issue in many visual tasks, as the output must respond to large enough areas in the image to capture information about large objects. We introduce the notion of an effective receptive field, and show that it both has a Gaussian distribution and only occupies a fraction of the full theoretical receptive field. We analyze the effective receptive field in several architecture designs, and the effect of nonlinear activations, dropout, sub-sampling and skip connections on it. This leads to suggestions for ways to address its tendency to be too small.
Cited by in corpus (74)
- MLP-Mixer: An all-MLP Architecture for Vision
- On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task
- Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing Images
- TransReID: Transformer-based Object Re-Identification
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- Video Frame Interpolation via Adaptive Separable Convolution
- Neural Network-Based Automatic Liver Tumor Segmentation With Random Forest-Based Candidate Filtering
- Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks
- LatentGNN: Learning Efficient Non-local Relations for Visual Recognition
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- SFD: Single Shot Scale-invariant Face Detector
- Mind the Pad -- CNNs can Develop Blind Spots
- CyCNN: A Rotation Invariant CNN using Polar Mapping and Cylindrical Convolution Layers
- A multi-task learning for cavitation detection and cavitation intensity recognition of valve acoustic signals
- Computation Reallocation for Object Detection
- Structured Attentions for Visual Question Answering
- Asymmetric 3D Context Fusion for Universal Lesion Detection
- Medical Image Segmentation Using Squeeze-and-Expansion Transformers
- Blurring the Line Between Structure and Learning to Optimize and Adapt Receptive Fields
- Global Aggregation then Local Distribution for Scene Parsing
- Improving Semantic Segmentation via Decoupled Body and Edge Supervision
- Spatial Dual-Modality Graph Reasoning for Key Information Extraction
- Funnel Activation for Visual Recognition
- Spatially-Attentive Patch-Hierarchical Network for Adaptive Motion Deblurring
- CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching
- Hyperspectral Image Classification With Context-Aware Dynamic Graph Convolutional Network
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Scale-aware Automatic Augmentation for Object Detection
- Low-Complexity Models for Acoustic Scene Classification Based on Receptive Field Regularization and Frequency Damping
- Geometric Style Transfer
- FeatherNets: Convolutional Neural Networks as Light as Feather for Face Anti-spoofing
- Positional Encoding as Spatial Inductive Bias in GANs
- Progressively Guided Alternate Refinement Network for RGB-D Salient Object Detection
- Domain-invariant Stereo Matching Networks
- Irregular Convolutional Neural Networks
- Boundary Aware U-Net for Glacier Segmentation
- Learning from Videos with Deep Convolutional LSTM Networks
- SegNBDT: Visual Decision Rules for Segmentation
- End-to-end Interpretable Neural Motion Planner
- Multi-Resolution Fully Convolutional Neural Networks for Monaural Audio Source Separation
- AFO-TAD: Anchor-free One-Stage Detector for Temporal Action Detection
- Deep Learning for Robust Motion Segmentation with Non-Static Cameras
- Representative Graph Neural Network
- Deep Convolutional Neural Network-based Bernoulli Heatmap for Head Pose Estimation
- Detecting Reflections by Combining Semantic and Instance Segmentation
- 3D Neighborhood Convolution: Learning Depth-Aware Features for RGB-D and RGB Semantic Segmentation
- Subjective and Objective De-raining Quality Assessment Towards Authentic Rain Image
- Learning To Pay Attention To Mistakes
- Visual Concept Reasoning Networks
- Recurrent Attention Model with Log-Polar Mapping is Robust against Adversarial Attacks
- GridTracer: Automatic Mapping of Power Grids using Deep Learning and Overhead Imagery
- Adaptive Context Network for Scene Parsing
- Bayesian deep learning for mapping via auxiliary information: a new era for geostatistics?
- Deformable Kernel Convolutional Network for Video Extreme Super-Resolution
- PGT: A Progressive Method for Training Models on Long Videos
- Self-supervised Learning with Fully Convolutional Networks
- A Spectral Nonlocal Block for Neural Networks
- Signature-Graph Networks
- GCCN: Global Context Convolutional Network
- Improving Skeleton-based Action Recognitionwith Robust Spatial and Temporal Features
- Extra Proximal-Gradient Inspired Non-local Network
- LUAI Challenge 2021 on Learning to Understand Aerial Images
- Enabling variable high spatial resolution retrieval from a long pulse BOTDA sensor
- Learnable Sampling 3D Convolution for Video Enhancement and Action Recognition
- Log-Polar Space Convolution for Convolutional Neural Networks
- Poly-NL: Linear Complexity Non-local Layers with Polynomials
- A Coarse-to-Fine Instance Segmentation Network with Learning Boundary Representation
- SAFE: Scale Aware Feature Encoder for Scene Text Recognition
- Scaling up deep neural networks: a capacity allocation perspective
- Learning Topology from Synthetic Data for Unsupervised Depth Completion
- Multi-Scale Convolutions for Learning Context Aware Feature Representations
- Density-embedding layers: a general framework for adaptive receptive fields
- Efficient Transfer Learning via Joint Adaptation of Network Architecture and Weight
- Encoder-Decoder based CNN and Fully Connected CRFs for Remote Sensed Image Segmentation