All you need is a good init
arXiv:1511.06422
Abstract
Layer-sequential unit-variance (LSUV) initialization - a simple method for weight initialization for deep net learning - is proposed. The method consists of the two steps. First, pre-initialize weights of each convolution or inner-product layer with orthonormal matrices. Second, proceed from the first to the final layer, normalizing the variance of the output of each layer to be equal to one. Experiment with different activation functions (maxout, ReLU-family, tanh) show that the proposed initialization leads to learning of very deep nets that (i) produces networks with test accuracy better or equal to standard methods and (ii) is at least as fast as the complex schemes proposed specifically for very deep nets such as FitNets (Romero et al. (2015)) and Highway (Srivastava et al. (2015)). Performance is evaluated on GoogLeNet, CaffeNet, FitNets and Residual nets and the state-of-the-art, or very close to it, is achieved on the MNIST, CIFAR-10/100 and ImageNet datasets.
Published as a conference paper at ICLR 2016
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Striving for Simplicity: The All Convolutional Net
- Going Deeper with Convolutions
- Spatially-sparse convolutional neural networks
- Random Walk Initialization for Training Very Deep Feedforward Networks
Cited by in corpus (85)
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Residual Networks of Residual Networks: Multilevel Residual Networks
- Systematic evaluation of CNN advances on the ImageNet
- Learning Constitutive Relations from Indirect Observations Using Deep Neural Networks
- ReZero is All You Need: Fast Convergence at Large Depth
- DSD: Dense-Sparse-Dense Training for Deep Neural Networks
- Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks
- An evidential classifier based on Dempster-Shafer theory and deep learning
- Optimization for deep learning: theory and algorithms
- N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning
- Fixup Initialization: Residual Learning Without Normalization
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
- Review: Deep Learning in Electron Microscopy
- Analysis and Optimization of Convolutional Neural Network Architectures
- Deep Hyperspherical Learning
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Colorization as a Proxy Task for Visual Understanding
- Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
- Revise Saturated Activation Functions
- Convolutional Sparse Kernel Network for Unsupervised Medical Image Analysis
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- Understanding Batch Normalization
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- A Comprehensive and Modularized Statistical Framework for Gradient Norm Equality in Deep Neural Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Learning Kernel for Conditional Moment-Matching Discrepancy-based Image Classification
- Conducting Credit Assignment by Aligning Local Representations
- M2CAI Workflow Challenge: Convolutional Neural Networks with Time Smoothing and Hidden Markov Model for Video Frames Classification
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Enhancing approximation abilities of neural networks by training derivatives
- Adjusting for Dropout Variance in Batch Normalization and Weight Initialization
- Knowledge Projection for Deep Neural Networks
- Optimization of Convolutional Neural Network using Microcanonical Annealing Algorithm
- Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?
- Improved Part Segmentation Performance by Optimising Realism of Synthetic Images using Cycle Generative Adversarial Networks
- Deep Residual Networks and Weight Initialization
- Convolutional Residual Memory Networks
- On Graph Classification Networks, Datasets and Baselines
- Norm-preserving Orthogonal Permutation Linear Unit Activation Functions (OPLU)
- Shifting Mean Activation Towards Zero with Bipolar Activation Functions
- NDT: Neual Decision Tree Towards Fully Functioned Neural Graph
- Normalization of Neural Networks using Analytic Variance Propagation
- Transfer entropy-based feedback improves performance in artificial neural networks
- A Scalable Approach for Facial Action Unit Classifier Training UsingNoisy Data for Pre-Training
- Neuron Campaign for Initialization Guided by Information Bottleneck Theory
- Activation function impact on Sparse Neural Networks
- Convolution Aware Initialization
- Supervised COSMOS Autoencoder: Learning Beyond the Euclidean Loss!
- Bridging the Gap Between Neural Networks and Neuromorphic Hardware with A Neural Network Compiler
- Deep Learning for Inverse Problems: Bounds and Regularizers
- Large-Scale Gradient-Free Deep Learning with Recursive Local Representation Alignment
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- Dynamic Sparse Graph for Efficient Deep Learning
- Seq2Tens: An Efficient Representation of Sequences by Low-Rank Tensor Projections
- Learning Less-Overlapping Representations
- A Survey of Techniques All Classifiers Can Learn from Deep Networks: Models, Optimizations, and Regularization
- Confined Gradient Descent: Privacy-preserving Optimization for Federated Learning
- Improving training of deep neural networks via Singular Value Bounding
- Learn Faster and Forget Slower via Fast and Stable Task Adaptation
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- An unsupervised long short-term memory neural network for event detection in cell videos
- Critical Percolation as a Framework to Analyze the Training of Deep Networks
- A Fully Trainable Network with RNN-based Pooling
- Sample Variance Decay in Randomly Initialized ReLU Networks
- Connecting Graph Convolutional Networks and Graph-Regularized PCA
- Classifying Textual Data with Pre-trained Vision Models through Transfer Learning and Data Transformations
- Zero Initialization of modified Gated Recurrent Encoder-Decoder Network for Short Term Load Forecasting
- Generalized Batch Normalization: Towards Accelerating Deep Neural Networks
- Data-driven Weight Initialization with Sylvester Solvers
- Orthogonal and Idempotent Transformations for Learning Deep Neural Networks
- Ambient Sound Provides Supervision for Visual Learning
- Light Multi-segment Activation for Model Compression
- WeightAlign: Normalizing Activations by Weight Alignment
- Self-Referenced Deep Learning
- A stepped sampling method for video detection using LSTM
- Optimizing Neural Network for Computer Vision task in Edge Device
- Deep Algorithms: designs for networks
- Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
- Learning to Initialize Gradient Descent Using Gradient Descent
- New Interpretations of Normalization Methods in Deep Learning
- Farkas layers: don't shift the data, fix the geometry
- Stochastic Sparse Learning with Momentum Adaptation for Imprecise Memristor Networks
- An Effective Training Method For Deep Convolutional Neural Network
- On the Modeling of Error Functions as High Dimensional Landscapes for Weight Initialization in Learning Networks