Prevalence of Neural Collapse during the terminal phase of deep learning training
arXiv:2008.08186 · doi:10.1073/pnas.2015509117
Abstract
Modern practice for training classification deepnets involves a Terminal Phase of Training (TPT), which begins at the epoch where training error first vanishes; During TPT, the training error stays effectively zero while training loss is pushed towards zero. Direct measurements of TPT, for three prototypical deepnet architectures and across seven canonical classification datasets, expose a pervasive inductive bias we call Neural Collapse, involving four deeply interconnected phenomena: (NC1) Cross-example within-class variability of last-layer training activations collapses to zero, as the individual activations themselves collapse to their class-means; (NC2) The class-means collapse to the vertices of a Simplex Equiangular Tight Frame (ETF); (NC3) Up to rescaling, the last-layer classifiers collapse to the class-means, or in other words to the Simplex ETF, i.e. to a self-dual configuration; (NC4) For a given activation, the classifier's decision collapses to simply choosing whichever class has the closest train class-mean, i.e. the Nearest Class Center (NCC) decision rule. The symmetric and very simple geometry induced by the TPT confers important benefits, including better generalization performance, better robustness, and better interpretability.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Memory-Efficient Implementation of DenseNets
- Convolutional Neural Networks Analyzed via Convolutional Sparse Coding
- Identifying Mislabeled Instances in Classification Datasets
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Deep Network Classification by Scattering and Homotopy Dictionary Learning
- When and How Can Deep Generative Models be Inverted?
Cited by in corpus (12)
- Mathematical Models of Overparameterized Neural Networks
- Neural Collapse with Cross-Entropy Loss
- Explicit regularization and implicit bias in deep network classifiers trained with the square loss
- Learning to Assimilate in Chaotic Dynamical Systems
- A Recipe for Global Convergence Guarantee in Deep Neural Networks
- Separation and Concentration in Deep Networks
- Distribution of Classification Margins: Are All Data Equal?
- Incremental Learning via Rate Reduction
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations
- Probing neural networks with t-SNE, class-specific projections and a guided tour
- Learning to Give Checkable Answers with Prover-Verifier Games
- Exploring the high dimensional geometry of HSI features