Gradient Starvation: A Learning Proclivity in Neural Networks
arXiv:2011.09468
Abstract
We identify and formalize a fundamental gradient descent phenomenon resulting in a learning proclivity in over-parameterized neural networks. Gradient Starvation arises when cross-entropy loss is minimized by capturing only a subset of features relevant for the task, despite the presence of other predictive features that fail to be discovered. This work provides a theoretical explanation for the emergence of such feature imbalance in neural networks. Using tools from Dynamical Systems theory, we identify simple properties of learning dynamics during gradient descent that lead to this imbalance, and prove that such a situation can be expected given certain statistical structure in training data. Based on our proposed formalism, we develop guarantees for a novel regularization method aimed at decoupling feature learning dynamics, improving accuracy and robustness in cases hindered by gradient starvation. We illustrate our findings with simple and real-world out-of-distribution (OOD) generalization experiments.
Proceeding of NeurIPS 2021
References in corpus (19)
- Shortcut Learning in Deep Neural Networks
- Understanding deep learning requires rethinking generalization
- A Closer Look at Memorization in Deep Networks
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- The Pitfalls of Simplicity Bias in Neural Networks
- Learning explanations that are hard to vary
- Learning from Failure: Training Debiased Classifier from Biased Classifier
- Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
- What shapes feature representations? Exploring datasets, architectures, and training
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- SGD on Neural Networks Learns Functions of Increasing Complexity
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Invariant Risk Minimization Games
- Cross-Entropy Loss and Low-Rank Features Have Responsibility for Adversarial Examples
- Implicit Regularization via Neural Feature Alignment
- Linear Regression Games: Convergence Guarantees to Approximate Out-of-Distribution Solutions
Cited by in corpus (11)
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
- Fairness via Representation Neutralization
- SAND-mask: An Enhanced Gradient Masking Strategy for the Discovery of Invariances in Domain Generalization
- Can Subnetwork Structure be the Key to Out-of-Distribution Generalization?
- The Low-Rank Simplicity Bias in Deep Networks
- Simple data balancing achieves competitive worst-group-accuracy
- Towards Understanding the Data Dependency of Mixup-style Training
- Quantifying and Improving Transferability in Domain Generalization
- Encouraging Intra-Class Diversity Through a Reverse Contrastive Loss for Better Single-Source Domain Generalization
- Combating Unknown Bias with Effective Bias-Conflicting Scoring and Gradient Alignment
- Counterfactual Supervision-based Information Bottleneck for Out-of-Distribution Generalization