Critical Learning Periods in Deep Neural Networks
arXiv:1711.08856
Abstract
Similar to humans and animals, deep artificial neural networks exhibit critical periods during which a temporary stimulus deficit can impair the development of a skill. The extent of the impairment depends on the onset and length of the deficit window, as in animal models, and on the size of the neural network. Deficits that do not affect low-level statistics, such as vertical flipping of the images, have no lasting effect on performance and can be overcome with further training. To better understand this phenomenon, we use the Fisher Information of the weights to measure the effective connectivity between layers of a network during training. Counterintuitively, information rises rapidly in the early phases of training, and then decreases, preventing redistribution of information resources in a phenomenon we refer to as a loss of "Information Plasticity". Our analysis suggests that the first few epochs are critical for the creation of strong connections that are optimal relative to the input data distribution. Once such strong connections are created, they do not appear to change during additional training. These findings suggest that the initial learning transient, under-scrutinized compared to asymptotic behavior, plays a key role in determining the outcome of the training process. Our findings, combined with recent theoretical results in the literature, also suggest that forgetting (decrease of information in the weights) is critical to achieving invariance and disentanglement in representation learning. Finally, critical periods are not restricted to biological systems, but can emerge naturally in learning systems, whether biological or artificial, due to fundamental constrains arising from learning dynamics and information processing.
References in corpus (1)
Cited by in corpus (28)
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- Triple Memory Networks: a Brain-Inspired Method for Continual Learning
- The Early Phase of Neural Network Training
- Data Augmentation Revisited: Rethinking the Distribution Gap between Clean and Augmented Data
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- Provably scale-covariant continuous hierarchical networks based on scale-normalized differential expressions coupled in cascade
- A Closer Look at Structured Pruning for Neural Network Compression
- LCA: Loss Change Allocation for Neural Network Training
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- ScaDLES: Scalable Deep Learning over Streaming data at the Edge
- DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression
- Smooth activations and reproducibility in deep networks
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- Anti-Distillation: Improving reproducibility of deep networks
- Accelerating Distributed ML Training via Selective Synchronization
- GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
- Nondeterminism and Instability in Neural Network Optimization
- Synthesizing Irreproducibility in Deep Networks
- Fast gradient-free activation maximization for neurons in spiking neural networks
- Flexible Communication for Optimal Distributed Learning over Unpredictable Networks
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- Intraclass clustering: an implicit learning ability that regularizes DNNs
- How many winning tickets are there in one DNN?
- CropDefender: deep watermark which is more convenient to train and more robust against cropping
- A Tale Of Two Long Tails
- Supervised Momentum Contrastive Learning for Few-Shot Classification
- A New MRAM-based Process In-Memory Accelerator for Efficient Neural Network Training with Floating Point Precision
- Domain Adaptor Networks for Hyperspectral Image Recognition