Normalization Techniques in Training DNNs: Methodology, Analysis and Application
arXiv:2009.12836
Abstract
Normalization techniques are essential for accelerating the training and improving the generalization of deep neural networks (DNNs), and have successfully been used in various applications. This paper reviews and comments on the past, present and future of normalization methods in the context of DNN training. We provide a unified picture of the main motivation behind different approaches from the perspective of optimization, and present a taxonomy for understanding the similarities and differences between them. Specifically, we decompose the pipeline of the most representative normalizing activation methods into three components: the normalization area partitioning, normalization operation and normalization representation recovery. In doing so, we provide insight for designing new normalization technique. Finally, we discuss the current progress in understanding normalization methods, and provide a comprehensive review of the applications of normalization for particular tasks, in which it can effectively solve the key issues.
20 pages
References in corpus (28)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Conditional Generative Adversarial Nets
- On the difficulty of training Recurrent Neural Networks
- A Learned Representation For Artistic Style
- Large Batch Training of Convolutional Networks
- Designing Neural Network Architectures using Reinforcement Learning
- L2 Regularization versus Batch and Weight Normalization
- Understanding and Improving Layer Normalization
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Orthogonal Weight Normalization: Solution to Optimization over Multiple Dependent Stiefel Manifolds in Deep Neural Networks
- Learning Feature Hierarchies with Centered Deep Boltzmann Machines
- Comparison of Batch Normalization and Weight Normalization Algorithms for the Large-scale Image Classification
- Batch Normalization is a Cause of Adversarial Vulnerability
- Normalizing the Normalizers: Comparing and Extending Network Normalization Schemes
- Generalized BackPropagation, Étude De Cas: Orthogonality
- Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks
- Optimization on Submanifolds of Convolution Kernels in CNNs
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform
- Streaming Normalization: Towards Simpler and More Biologically-plausible Normalizations for Online and Recurrent Learning
- DizzyRNN: Reparameterizing Recurrent Neural Networks for Norm-Preserving Backpropagation
- Projection Based Weight Normalization for Deep Neural Networks
- Low-Precision Batch-Normalized Activations
- Be Like Water: Robustness to Extraneous Variables Via Adaptive Feature Normalization
- Deep Learning for Inverse Problems: Bounds and Regularizers
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- Unpaired Image Translation via Adaptive Convolution-based Normalization
- Large Batch Training Does Not Need Warmup
- Orthogonal Wasserstein GANs