Convolutional Normalization: Improving Deep Convolutional Network Robustness and Training
arXiv:2103.00673
Abstract
Normalization techniques have become a basic component in modern convolutional neural networks (ConvNets). In particular, many recent works demonstrate that promoting the orthogonality of the weights helps train deep models and improve robustness. For ConvNets, most existing methods are based on penalizing or normalizing weight matrices derived from concatenating or flattening the convolutional kernels. These methods often destroy or ignore the benign convolutional structure of the kernels; therefore, they are often expensive or impractical for deep ConvNets. In contrast, we introduce a simple and efficient "Convolutional Normalization" (ConvNorm) method that can fully exploit the convolutional structure in the Fourier domain and serve as a simple plug-and-play module to be conveniently incorporated into any ConvNets. Our method is inspired by recent work on preconditioning methods for convolutional sparse coding and can effectively promote each layer's channel-wise isometry. Furthermore, we show that our ConvNorm can reduce the layerwise spectral norm of the weight matrices and hence improve the Lipschitzness of the network, leading to easier training and improved robustness for deep ConvNets. Applied to classification under noise corruptions and generative adversarial network (GAN), we show that the ConvNorm improves the robustness of common ConvNets such as ResNet and the performance of GAN. We verify our findings via numerical experiments on CIFAR and ImageNet.
SL and XL contributed equally to this work; 23 pages, 6 figures, 6 tables, published in NeurIPS'21
References in corpus (18)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Exploring Simple Siamese Representation Learning
- A Closer Look at Memorization in Deep Networks
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- Orthogonal Weight Normalization: Solution to Optimization over Multiple Dependent Stiefel Manifolds in Deep Neural Networks
- On Loss Functions for Deep Neural Networks in Classification
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Generalized BackPropagation, Étude De Cas: Orthogonality
- Lipschitz Generative Adversarial Nets
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform
- Deep Isometric Learning for Visual Recognition
- Orthogonalizing Convolutional Layers with the Cayley Transform
- Isometric Autoencoders
- Finding the Sparsest Vectors in a Subspace: Theory, Algorithms, and Applications
- Approximated Orthonormal Normalisation in Training Neural Networks