Making Convolutional Networks Shift-Invariant Again
arXiv:1904.11486
Abstract
Modern convolutional networks are not shift-invariant, as small input shifts or translations can cause drastic changes in the output. Commonly used downsampling methods, such as max-pooling, strided-convolution, and average-pooling, ignore the sampling theorem. The well-known signal processing fix is anti-aliasing by low-pass filtering before downsampling. However, simply inserting this module into deep networks degrades performance; as a result, it is seldomly used today. We show that when integrated correctly, it is compatible with existing architectural components, such as max-pooling and strided-convolution. We observe \textit{increased accuracy} in ImageNet classification, across several commonly-used architectures, such as ResNet, DenseNet, and MobileNet, indicating effective regularization. Furthermore, we observe \textit{better generalization}, in terms of stability and robustness to input corruptions. Our results demonstrate that this classical signal processing technique has been undeservingly overlooked in modern deep networks. Code and anti-aliased versions of popular networks are available at https://richzhang.github.io/antialiased-cnns/ .
Accepted to ICML 2019
Cited by in corpus (40)
- Score-Based Generative Modeling through Stochastic Differential Equations
- Alias-Free Generative Adversarial Networks
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis
- Measuring Robustness to Natural Distribution Shifts in Image Classification
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- 3D Scene Geometry Estimation from 360 Imagery: A Survey
- Mind the Pad -- CNNs can Develop Blind Spots
- Delving Deeper into Anti-aliasing in ConvNets
- Knee Injury Detection using MRI with Efficiently-Layered Network (ELNet)
- Does enhanced shape bias improve neural network robustness to common corruptions?
- Total Deep Variation: A Stable Regularizer for Inverse Problems
- DexPilot: Vision Based Teleoperation of Dexterous Robotic Hand-Arm System
- Coordinate Independent Convolutional Networks -- Isometry and Gauge Equivariant Convolutions on Riemannian Manifolds
- PIPAL: a Large-Scale Image Quality Assessment Dataset for Perceptual Image Restoration
- A survey on Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
- Transfering Low-Frequency Features for Domain Adaptation
- LiftPool: Bidirectional ConvNet Pooling
- Fre-GAN: Adversarial Frequency-consistent Audio Synthesis
- NeRD: Neural Representation of Distribution for Medical Image Segmentation
- Shift Equivariance in Object Detection
- PropagationNet: Propagate Points to Curve to Learn Structure Information
- Improving Sound Event Classification by Increasing Shift Invariance in Convolutional Neural Networks
- Anti-aliasing Deep Image Classifiers using Novel Depth Adaptive Blurring and Activation Function
- Why Accuracy Is Not Enough: The Need for Consistency in Object Detection
- Learning to Factorize and Relight a City
- Data augmentation and image understanding
- Learning Translation Invariance in CNNs
- UniNet: Unified Architecture Search with Convolution, Transformer, and MLP
- Revisiting Adaptive Convolutions for Video Frame Interpolation
- What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy Conditions
- Frequency Pooling: Shift-Equivalent and Anti-Aliasing Downsampling
- On the Frequency Bias of Generative Models
- Modeling the Nonsmoothness of Modern Neural Networks
- Learning Equivariant Representations
- Shape Adaptor: A Learnable Resizing Module
- Exploring Novel Pooling Strategies for Edge Preserved Feature Maps in Convolutional Neural Networks
- Group Equivariant Subsampling
- Neuron segmentation using 3D wavelet integrated encoder-decoder network
- How saccadic vision might help with theinterpretability of deep networks