Prevalence of Neural Collapse during the terminal phase of deep learning training
arXiv:2008.08186 · doi:10.1073/pnas.2015509117
Abstract
Modern practice for training classification deepnets involves a Terminal Phase of Training (TPT), which begins at the epoch where training error first vanishes; During TPT, the training error stays effectively zero while training loss is pushed towards zero. Direct measurements of TPT, for three prototypical deepnet architectures and across seven canonical classification datasets, expose a pervasive inductive bias we call Neural Collapse, involving four deeply interconnected phenomena: (NC1) Cross-example within-class variability of last-layer training activations collapses to zero, as the individual activations themselves collapse to their class-means; (NC2) The class-means collapse to the vertices of a Simplex Equiangular Tight Frame (ETF); (NC3) Up to rescaling, the last-layer classifiers collapse to the class-means, or in other words to the Simplex ETF, i.e. to a self-dual configuration; (NC4) For a given activation, the classifier's decision collapses to simply choosing whichever class has the closest train class-mean, i.e. the Nearest Class Center (NCC) decision rule. The symmetric and very simple geometry induced by the TPT confers important benefits, including better generalization performance, better robustness, and better interpretability.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Memory-Efficient Implementation of DenseNets
- Convolutional Neural Networks Analyzed via Convolutional Sparse Coding
- Identifying Mislabeled Instances in Classification Datasets
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Deep Network Classification by Scattering and Homotopy Dictionary Learning
- When and How Can Deep Generative Models be Inverted?
Cited by in corpus (34)
- A Survey on Hyperdimensional Computing aka Vector Symbolic Architectures, Part II: Applications, Cognitive Models, and Challenges
- Neural Collapse Inspired Attraction-Repulsion-Balanced Loss for Imbalanced Learning
- Mathematical Models of Overparameterized Neural Networks
- Understanding Deep Learning via Decision Boundary
- Closed-Loop Data Transcription to an LDR via Minimaxing Rate Reduction
- Gaussian Universality of Perceptrons with Random Labels
- Prototypical Partial Optimal Transport for Universal Domain Adaptation
- Neural Collapse with Cross-Entropy Loss
- Federated Optimization of Smooth Loss Functions
- An Unconstrained Layer-Peeled Perspective on Neural Collapse
- Coding schemes in neural networks learning classification tasks
- Explicit regularization and implicit bias in deep network classifiers trained with the square loss
- A Recipe for Global Convergence Guarantee in Deep Neural Networks
- Separation and Concentration in Deep Networks
- Learning to Assimilate in Chaotic Dynamical Systems
- CoReS: Compatible Representations via Stationarity
- Learn by Reasoning: Analogical Weight Generation for Few-Shot Class-Incremental Learning
- Theoretical Guarantees for Low-Rank Compression of Deep Neural Networks
- Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
- Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
- Distribution of Classification Margins: Are All Data Equal?
- Incremental Learning via Rate Reduction
- Probing neural networks with t-SNE, class-specific projections and a guided tour
- Evaluating the Fairness of Neural Collapse in Medical Image Classification
- Double descent in quantum kernel methods
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations
- Dependence Induced Representations
- Reviewing Model Collapse and Countermeasures
- Deep Minimax Classifiers for Imbalanced Datasets with a Small Number of Minority Samples
- Gradient flow in parameter space is equivalent to linear interpolation in output space
- Unsupervised Learning of Unbiased Visual Representations
- Learning to Give Checkable Answers with Prover-Verifier Games
- Exploring the high dimensional geometry of HSI features
- Rethinking the Backbone in Class Imbalanced Federated Source Free Domain Adaptation: The Utility of Vision Foundation Models