Data-dependent Initializations of Convolutional Neural Networks
arXiv:1511.06856
Abstract
Convolutional Neural Networks spread through computer vision like a wildfire, impacting almost all visual tasks imaginable. Despite this, few researchers dare to train their models from scratch. Most work builds on one of a handful of ImageNet pre-trained models, and fine-tunes or adapts these for specific tasks. This is in large part due to the difficulty of properly initializing these networks from scratch. A small miscalibration of the initial weights leads to vanishing or exploding gradients, as well as poor convergence properties. In this work we present a fast and simple data-dependent initialization procedure, that sets the weights of a network such that all units in the network train at roughly the same rate, avoiding vanishing or exploding gradients. Our initialization matches the current state-of-the-art unsupervised or self-supervised pre-training methods on standard computer vision tasks, such as image classification and object detection, while being roughly three orders of magnitude faster. When combined with pre-training methods, our initialization significantly outperforms prior work, narrowing the gap between supervised and unsupervised pre-training.
ICLR 2016
References in corpus (4)
Cited by in corpus (61)
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- Unsupervised Representation Learning by Predicting Image Rotations
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Adversarial Feature Learning
- Contrastive Multiview Coding
- Deep Clustering for Unsupervised Learning of Visual Features
- Regularization for Deep Learning: A Taxonomy
- Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination
- Unsupervised Learning via Meta-Learning
- Improvements to context based self-supervised learning
- Streaming convolutional neural networks for end-to-end learning with multi-megapixel images
- A critical analysis of self-supervision, or what we can learn from a single image
- Local Aggregation for Unsupervised Learning of Visual Embeddings
- Unsupervised Meta-Learning for Reinforcement Learning
- Colorization as a Proxy Task for Visual Understanding
- PCA-Initialized Deep Neural Networks Applied To Document Image Analysis
- AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations rather than Data
- Optimization on Submanifolds of Convolution Kernels in CNNs
- Look, Listen and Learn
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Colorful Image Colorization
- Unsupervised Representation Learning by Sorting Sequences
- Learning Image Representations by Completing Damaged Jigsaw Puzzles
- Historical Document Image Segmentation with LDA-Initialized Deep Neural Networks
- BYOL works even without batch statistics
- Self-supervised learning of visual features through embedding images into text topic spaces
- Weight Initialization of Deep Neural Networks(DNNs) using Data Statistics
- Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction
- Privacy Risk in Machine Learning: Analyzing the Connection to Overfitting
- Gradients as Features for Deep Representation Learning
- Norm-preserving Orthogonal Permutation Linear Unit Activation Functions (OPLU)
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Neuron Campaign for Initialization Guided by Information Bottleneck Theory
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
- AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- Dynamical Isometry: The Missing Ingredient for Neural Network Pruning
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- TextTopicNet - Self-Supervised Learning of Visual Features Through Embedding Images on Semantic Text Spaces
- Self-Supervised Visual Representations for Cross-Modal Retrieval
- Random Projection in Deep Neural Networks
- An unsupervised long short-term memory neural network for event detection in cell videos
- Auto-tuning of Deep Neural Networks by Conflicting Layer Removal
- Sample Variance Decay in Randomly Initialized ReLU Networks
- AETv2: AutoEncoding Transformations for Self-Supervised Representation Learning by Minimizing Geodesic Distances in Lie Groups
- Laplacian Denoising Autoencoder
- Connecting Graph Convolutional Networks and Graph-Regularized PCA
- Data-driven Weight Initialization with Sylvester Solvers
- Self-supervised Learning with Fully Convolutional Networks
- Gabor filter incorporated CNN for compression
- Collaboration among Image and Object Level Features for Image Colourisation
- Self-Referenced Deep Learning
- Ambient Sound Provides Supervision for Visual Learning
- Multiclass non-Adversarial Image Synthesis, with Application to Classification from Very Small Sample
- Improving Visual Recognition using Ambient Sound for Supervision
- Accelerated Dual Learning by Homotopic Initialization