Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
arXiv:1602.07868
Abstract
We present weight normalization: a reparameterization of the weight vectors in a neural network that decouples the length of those weight vectors from their direction. By reparameterizing the weights in this way we improve the conditioning of the optimization problem and we speed up convergence of stochastic gradient descent. Our reparameterization is inspired by batch normalization but does not introduce any dependencies between the examples in a minibatch. This means that our method can also be applied successfully to recurrent models such as LSTMs and to noise-sensitive applications such as deep reinforcement learning or generative models, for which batch normalization is less well suited. Although our method is much simpler, it still provides much of the speed-up of full batch normalization. In addition, the computational overhead of our method is lower, permitting more optimization steps to be taken in the same amount of time. We demonstrate the usefulness of our method on applications in supervised image recognition, generative modelling, and deep reinforcement learning.
References in corpus (6)
Cited by in corpus (155)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Spectral Normalization for Generative Adversarial Networks
- Progressive Growing of GANs for Improved Quality, Stability, and Variation
- Temporal Ensembling for Semi-Supervised Learning
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- NiftyNet: a deep-learning platform for medical imaging
- Recurrence is required to capture the representational dynamics of the human visual system
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Learning to Generate Reviews and Discovering Sentiment
- Deep Appearance Models for Face Rendering
- GAN-Leaks: A Taxonomy of Membership Inference Attacks against Generative Models
- Neural Lander: Stable Drone Landing Control using Learned Dynamics
- PhyCRNet: Physics-informed Convolutional-Recurrent Network for Solving Spatiotemporal PDEs
- Robust Large Margin Deep Neural Networks
- Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models
- Neuromorphic Deep Learning Machines
- Online and Linear-Time Attention by Enforcing Monotonic Alignments
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- L1-Norm Batch Normalization for Efficient Training of Deep Neural Networks
- Pythia v0.1: the Winning Entry to the VQA Challenge 2018
- Differentiable Learning-to-Normalize via Switchable Normalization
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- Super Resolution Convolutional Neural Network Models for Enhancing Resolution of Rock Micro-CT Images
- Sharp Minima Can Generalize For Deep Nets
- Compression-aware Training of Deep Networks
- DTAAD: Dual Tcn-Attention Networks for Anomaly Detection in Multivariate Time Series Data
- A survey on GANs for computer vision: Recent research, analysis and taxonomy
- Fixup Initialization: Residual Learning Without Normalization
- Streaming convolutional neural networks for end-to-end learning with multi-megapixel images
- Review: Deep Learning in Electron Microscopy
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- Adversarial Attacks Against Deep Generative Models on Data: A Survey
- DIVA: Domain Invariant Variational Autoencoders
- Multiplicative LSTM for sequence modelling
- Balanced Sparsity for Efficient DNN Inference on GPU
- Bilinear Attention Networks
- DiracNets: Training Very Deep Neural Networks Without Skip-Connections
- Theoretical Analysis of Auto Rate-Tuning by Batch Normalization
- Deep Reinforcement Learning Control of Quantum Cartpoles
- On the Effects of Batch and Weight Normalization in Generative Adversarial Networks
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Dynamic Evaluation of Neural Sequence Models
- Comparison of Batch Normalization and Weight Normalization Algorithms for the Large-scale Image Classification
- Learning Algorithms for Active Learning
- Convolutional Neural Networks for Continuous QoE Prediction in Video Streaming Services
- Convolutional Sparse Kernel Network for Unsupervised Medical Image Analysis
- Learning Continuous Semantic Representations of Symbolic Expressions
- Attention U-Net as a surrogate model for groundwater prediction
- Improving Electron Micrograph Signal-to-Noise with an Atrous Convolutional Encoder-Decoder
- Understanding Batch Normalization
- Optimization on Submanifolds of Convolution Kernels in CNNs
- Novelty Detection with GAN
- Second-order Convolutional Neural Networks
- Online Normalization for Training Neural Networks
- A Comprehensive and Modularized Statistical Framework for Gradient Norm Equality in Deep Neural Networks
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- Safer Classification by Synthesis
- Few-Shot Generalization Across Dialogue Tasks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Self-Supervised Visual Learning by Variable Playback Speeds Prediction of a Video
- A Flow-Based Neural Network for Time Domain Speech Enhancement
- Who Needs Words? Lexicon-Free Speech Recognition
- Amortized Variational Inference: A Systematic Review
- Entropic gradient descent algorithms and wide flat minima
- Very Lightweight Photo Retouching Network with Conditional Sequential Modulation
- Direction Concentration Learning: Enhancing Congruency in Machine Learning
- Block-Matching Convolutional Neural Network for Image Denoising
- Smooth Neighbors on Teacher Graphs for Semi-supervised Learning
- Comparing Normalization Methods for Limited Batch Size Segmentation Neural Networks
- Product-based Neural Networks for User Response Prediction over Multi-field Categorical Data
- Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches
- Self-Supervised Variational Auto-Encoders
- Geometric Approaches to Increase the Expressivity of Deep Neural Networks for MR Reconstruction
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Normalized Direction-preserving Adam
- Here comes the SU(N): multivariate quantum gates and gradients
- Comparison of Deep learning models on time series forecasting : a case study of Dissolved Oxygen Prediction
- A Deep Generative Model of Speech Complex Spectrograms
- Unsupervised Training for 3D Morphable Model Regression
- Semi-supervised Rare Disease Detection Using Generative Adversarial Network
- VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions
- Score and Lyrics-Free Singing Voice Generation
- Adjusting for Dropout Variance in Batch Normalization and Weight Initialization
- Transforming task representations to perform novel tasks
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Revisit Fuzzy Neural Network: Demystifying Batch Normalization and ReLU with Generalized Hamming Network
- The Benefits of Over-parameterization at Initialization in Deep ReLU Networks
- Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?
- Projection Based Weight Normalization for Deep Neural Networks
- Riemannian approach to batch normalization
- Optimization Theory for ReLU Neural Networks Trained with Normalization Layers
- LOGAN: Membership Inference Attacks Against Generative Models
- SSFL: Tackling Label Deficiency in Federated Learning via Personalized Self-Supervision
- Fix your classifier: the marginal value of training the last weight layer
- Viewpoint-Aware Loss with Angular Regularization for Person Re-Identification
- Generating Neural Networks with Neural Networks
- GraN-GAN: Piecewise Gradient Normalization for Generative Adversarial Networks
- Ring loss: Convex Feature Normalization for Face Recognition
- Regularizing by the Variance of the Activations' Sample-Variances
- SSN: Learning Sparse Switchable Normalization via SparsestMax
- Implicit Rugosity Regularization via Data Augmentation
- Physics-Informed Neural Networks for Transonic Flows around an Airfoil
- Statistically Motivated Second Order Pooling
- A computational framework for nanotrusses: input convex neural networks approach
- Accelerating Natural Gradient with Higher-Order Invariance
- PixelCNN Models with Auxiliary Variables for Natural Image Modeling
- Sequence to Point Learning Based on Bidirectional Dilated Residual Network for Non Intrusive Load Monitoring
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Making CNNs for Video Parsing Accessible
- Neural Networks with Small Weights and Depth-Separation Barriers
- Softmax Dissection: Towards Understanding Intra- and Inter-class Objective for Embedding Learning
- Training Deep Neural Networks Without Batch Normalization
- A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent
- Batch-normalized Recurrent Highway Networks
- Blind Source Separation of Single-Channel Mixtures via Multi-Encoder Autoencoders
- toon2real: Translating Cartoon Images to Realistic Images
- Catch-A-Waveform: Learning to Generate Audio from a Single Short Example
- Adaptive Precision Training (AdaPT): A dynamic fixed point quantized training approach for DNNs
- Proportionate gradient updates with PercentDelta
- Simplified Stochastic Feedforward Neural Networks
- Normalization in Training U-Net for 2D Biomedical Semantic Segmentation
- Mean Shift Rejection: Training Deep Neural Networks Without Minibatch Statistics or Normalization
- Improving training of deep neural networks via Singular Value Bounding
- Instance Enhancement Batch Normalization: an Adaptive Regulator of Batch Noise
- Discriminative Multi-modality Speech Recognition
- Regularity Normalization: Neuroscience-Inspired Unsupervised Attention across Neural Network Layers
- Homocentric Hypersphere Feature Embedding for Person Re-identification
- Investigating and Improving Latent Density Segmentation Models for Aleatoric Uncertainty Quantification in Medical Imaging
- Data-Adaptive Discriminative Feature Localization with Statistically Guaranteed Interpretation
- Model-agnostic out-of-distribution detection using combined statistical tests
- Recurrent Point Review Models
- A Neural Network model with Bidirectional Whitening
- FixNorm: Dissecting Weight Decay for Training Deep Neural Networks
- Approximated Orthonormal Normalisation in Training Neural Networks
- Understanding the Disharmony between Weight Normalization Family and Weight Decay: shifted Regularizer
- Generalized Batch Normalization: Towards Accelerating Deep Neural Networks
- On Batch Orthogonalization Layers
- Deep Control - a simple automatic gain control for memory efficient and high performance training of deep convolutional neural networks
- Double Forward Propagation for Memorized Batch Normalization
- Logographic Subword Model for Neural Machine Translation
- An Informal Introduction to Multiplet Neural Networks
- High Diversity Attribute Guided Face Generation with GANs
- An Effective Training Method For Deep Convolutional Neural Network
- Modeling Grasp Motor Imagery through Deep Conditional Generative Models
- Optimization on Product Submanifolds of Convolution Kernels
- Learning Inward Scaled Hypersphere Embedding: Exploring Projections in Higher Dimensions
- Efficient Modelling Across Time of Human Actions and Interactions
- Neighbourhood Distillation: On the benefits of non end-to-end distillation
- A scale invariant ranking function for learning-to-rank: a real-world use case
- RotationOut as a Regularization Method for Neural Network
- Diversity Regularized Adversarial Learning
- Approximate Random Dropout