Are All Layers Created Equal?
arXiv:1902.01996
Abstract
Understanding deep neural networks is a major research objective with notable experimental and theoretical attention in recent years. The practical success of excessively large networks underscores the need for better theoretical analyses and justifications. In this paper we focus on layer-wise functional structure and behavior in overparameterized deep models. To do so, we study empirically the layers' robustness to post-training re-initialization and re-randomization of the parameters. We provide experimental results which give evidence for the heterogeneity of layers. Morally, layers of large deep neural networks can be categorized as either "robust" or "critical". Resetting the robust layers to their initial values does not result in adverse decline in performance. In many cases, robust layers hardly change throughout training. In contrast, re-initializing critical layers vastly degrades the performance of the network with test error essentially dropping to random guesses. Our study provides further evidence that mere parameter counting or norm calculations are too coarse in studying generalization of deep models, and "flatness" and robustness analysis of trained models need to be examined while taking into account the respective network architectures.
JMLR 2022, 28 pages, 21 figures
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Explaining and Harnessing Adversarial Examples
- A Convergence Theory for Deep Learning via Over-Parameterization
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Small ReLU networks are powerful memorizers: a tight analysis of memorization capacity
- Identifying Generalization Properties in Neural Networks
- Deep vs. shallow networks : An approximation theory perspective
Cited by in corpus (44)
- Transfusion: Understanding Transfer Learning for Medical Imaging
- What is being transferred in transfer learning?
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- The Modern Mathematics of Deep Learning
- Interpreting and Improving Adversarial Robustness of Deep Neural Networks with Neuron Sensitivity
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
- ReLeQ: A Reinforcement Learning Approach for Deep Quantization of Neural Networks
- Predicting Neural Network Accuracy from Weights
- When BERT Plays the Lottery, All Tickets Are Winning
- How fine can fine-tuning be? Learning efficient language models
- Visualizing and Understanding the Effectiveness of BERT
- Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
- Object Detector Differences when using Synthetic and Real Training Data
- LCA: Loss Change Allocation for Neural Network Training
- A Broader Study of Cross-Domain Few-Shot Learning
- PAC-Bayes with Backprop
- Radio source-component association for the LOFAR Two-metre Sky Survey with region-based convolutional neural networks
- Training Robust Deep Neural Networks via Adversarial Noise Propagation
- The intriguing role of module criticality in the generalization of deep networks
- On the geometry of generalization and memorization in deep neural networks
- Taylorized Training: Towards Better Approximation of Neural Network Training at Finite Width
- ReFine: Re-randomization before Fine-tuning for Cross-domain Few-shot Learning
- Efficient and Private Federated Learning with Partially Trainable Networks
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Learning PAC-Bayes Priors for Probabilistic Neural Networks
- HourNAS: Extremely Fast Neural Architecture Search Through an Hourglass Lens
- Diet deep generative audio models with structured lottery
- Compare Where It Matters: Using Layer-Wise Regularization To Improve Federated Learning on Heterogeneous Data
- Fast Hardware-Aware Neural Architecture Search
- Hyperplane Arrangements of Trained ConvNets Are Biased
- Understanding Learning Dynamics for Neural Machine Translation
- Experiments with Rich Regime Training for Deep Learning
- Rethinking the Value of Transformer Components
- E2-Train: Training State-of-the-art CNNs with Over 80% Energy Savings
- Grassmannian Packings in Neural Networks: Learning with Maximal Subspace Packings for Diversity and Anti-Sparsity
- Properties of the After Kernel
- Multirate Training of Neural Networks
- Self-Teaching Networks
- What can linear interpolation of neural network loss landscapes tell us?
- On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks
- Neighbourhood Distillation: On the benefits of non end-to-end distillation
- Identifying Layers Susceptible to Adversarial Attacks