Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
arXiv:1805.12076
Abstract
Despite existing work on ensuring generalization of neural networks in terms of scale sensitive complexity measures, such as norms, margin and sharpness, these complexity measures do not offer an explanation of why neural networks generalize better with over-parametrization. In this work we suggest a novel complexity measure based on unit-wise capacities resulting in a tighter generalization bound for two layer ReLU networks. Our capacity bound correlates with the behavior of test error with increasing network sizes, and could potentially explain the improvement in generalization with over-parametrization. We further present a matching lower bound for the Rademacher complexity that improves over previous capacity lower bounds for neural networks.
19 pages, 8 figures
References in corpus (6)
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Spectrally-normalized margin bounds for neural networks
- Fisher-Rao Metric, Geometry, and Complexity of Neural Networks
- Stronger generalization bounds for deep nets via a compression approach
- Norm-Based Capacity Control in Neural Networks
Cited by in corpus (103)
- A Survey on Visual Transformer
- Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
- Theory of overparametrization in quantum neural networks
- Deep learning observables in computational fluid dynamics
- Scaling description of generalization with number of parameters in deep learning
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- Intrinsic dimension of data representations in deep neural networks
- Fantastic Generalization Measures and Where to Find Them
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Understanding Dimensional Collapse in Contrastive Self-supervised Learning
- Continual Learning via Neural Pruning
- The Global Landscape of Neural Networks: An Overview
- Analyzing Upper Bounds on Mean Absolute Errors for Deep Neural Network Based Vector-to-Vector Regression
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- The Early Phase of Neural Network Training
- Winning the Lottery with Continuous Sparsification
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- Identifying Mislabeled Data using the Area Under the Margin Ranking
- What Makes Multi-modal Learning Better than Single (Provably)
- Finite Versus Infinite Neural Networks: an Empirical Study
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- Error estimates of residual minimization using neural networks for linear PDEs
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Lost in Pruning: The Effects of Pruning Neural Networks beyond Test Accuracy
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- The Implicit and Explicit Regularization Effects of Dropout
- Greedy Layerwise Learning Can Scale to ImageNet
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- How Does Mixup Help With Robustness and Generalization?
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
- How does Weight Correlation Affect the Generalisation Ability of Deep Neural Networks
- Data-Independent Neural Pruning via Coresets
- Network Pruning That Matters: A Case Study on Retraining Variants
- Observational Overfitting in Reinforcement Learning
- On Power Laws in Deep Ensembles
- The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
- The Benefits of Over-parameterization at Initialization in Deep ReLU Networks
- Mean-Field Neural ODEs via Relaxed Optimal Control
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Understanding the role of importance weighting for deep learning
- Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima
- The intriguing role of module criticality in the generalization of deep networks
- Distributional Generalization: A New Kind of Generalization
- Sanity-Checking Pruning Methods: Random Tickets can Win the Jackpot
- Dropout: Explicit Forms and Capacity Control
- Generalization bounds for deep learning
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
- Vector Contraction for Rademacher Complexity
- Towards Learning Convolutions from Scratch
- Efficient and Private Federated Learning with Partially Trainable Networks
- Understanding Generalization in Deep Learning via Tensor Methods
- Towards Task and Architecture-Independent Generalization Gap Predictors
- Convex Geometry and Duality of Over-parameterized Neural Networks
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- An Exponential Improvement on the Memorization Capacity of Deep Threshold Networks
- Deep Learning is Singular, and That's Good
- For self-supervised learning, Rationality implies generalization, provably
- Approximation power of random neural networks
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- On the interplay between data structure and loss function in classification problems
- Policy Search with Rare Significant Events: Choosing the Right Partner to Cooperate with
- Refactoring Neural Networks for Verification
- Stochastic Training is Not Necessary for Generalization
- Distance-Based Regularisation of Deep Networks for Fine-Tuning
- Implicit Regularization via Neural Feature Alignment
- Deep Learning for Inverse Problems: Bounds and Regularizers
- Measuring Generalization with Optimal Transport
- A Theoretical Analysis of Fine-tuning with Linear Teachers
- Weak and Strong Gradient Directions: Explaining Memorization, Generalization, and Hardness of Examples at Scale
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Provably Efficient Neural Estimation of Structural Equation Model: An Adversarial Approach
- Empirical Risk Minimization in the Interpolating Regime with Application to Neural Network Learning
- Data-Independent Structured Pruning of Neural Networks via Coresets
- Deep learning for pedestrians: backpropagation in CNNs
- Adversarial Risk Bounds for Neural Networks through Sparsity based Compression
- Identifying Critical Neurons in ANN Architectures using Mixed Integer Programming
- Using Wavelets and Spectral Methods to Study Patterns in Image-Classification Datasets
- Tangent Space Separability in Feedforward Neural Networks
- Improved generalization by noise enhancement
- Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
- Abstraction Mechanisms Predict Generalization in Deep Neural Networks
- On Predicting Generalization using GANs
- Generalization Error of Generalized Linear Models in High Dimensions
- Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
- Consensus-based Interpretable Deep Neural Networks with Application to Mortality Prediction
- Achieving Small Test Error in Mildly Overparameterized Neural Networks
- Provably Training Overparameterized Neural Network Classifiers with Non-convex Constraints
- Global Convergence and Geometric Characterization of Slow to Fast Weight Evolution in Neural Network Training for Classifying Linearly Non-Separable Data
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Fairness Sample Complexity and the Case for Human Intervention
- Dimension Independent Generalization Error by Stochastic Gradient Descent
- The role of a layer in deep neural networks: a Gaussian Process perspective
- Tangent Space Sensitivity and Distribution of Linear Regions in ReLU Networks
- On the generalization of bayesian deep nets for multi-class classification
- Characterization of Excess Risk for Locally Strongly Convex Population Risk
- Risk Bounds for Learning via Hilbert Coresets
- CNN with large memory layers
- Model-Aware Regularization For Learning Approaches To Inverse Problems
- How isotropic kernels perform on simple invariants
- Mean-field Analysis of Piecewise Linear Solutions for Wide ReLU Networks