Spectrally-normalized margin bounds for neural networks
arXiv:1706.08498
Abstract
This paper presents a margin-based multiclass generalization bound for neural networks that scales with their margin-normalized "spectral complexity": their Lipschitz constant, meaning the product of the spectral norms of the weight matrices, times a certain correction factor. This bound is empirically investigated for a standard AlexNet network trained with SGD on the mnist and cifar10 datasets, with both original and random labels; the bound, the Lipschitz constants, and the excess risks are all in direct correlation, suggesting both that SGD selects predictors whose complexity scales with the difficulty of the learning task, and secondly that the presented bound is sensitive to this complexity.
Comparison to arXiv v1: 1-norm in main bound refined to (2,1)-group-norm. Comparison to NIPS camera ready: typo fixes
References in corpus (2)
Cited by in corpus (26)
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- The no-free-lunch theorems of supervised learning
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Implicit Regularization in Deep Learning
- What Makes Multi-modal Learning Better than Single (Provably)
- SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Minimax Estimation of Conditional Moment Models
- Generalization Bounds for Convolutional Neural Networks
- Generalization Bounds For Unsupervised and Semi-Supervised Learning With Autoencoders
- A PAC-Bayesian Approach to Generalization Bounds for Graph Neural Networks
- Robustness to Pruning Predicts Generalization in Deep Neural Networks
- On the Role of Dataset Quality and Heterogeneity in Model Confidence
- Markov-Lipschitz Deep Learning
- Asymptotic Singular Value Distribution of Linear Convolutional Layers
- Noether: The More Things Change, the More Stay the Same
- Generalization bounds for graph convolutional neural networks via Rademacher complexity
- Single Layer Predictive Normalized Maximum Likelihood for Out-of-Distribution Detection
- Compressing Heavy-Tailed Weight Matrices for Non-Vacuous Generalization Bounds
- Inferential Wasserstein Generative Adversarial Networks
- Achieving Small Test Error in Mildly Overparameterized Neural Networks
- Revisiting hard thresholding for DNN pruning
- From deep to Shallow: Equivalent Forms of Deep Networks in Reproducing Kernel Krein Space and Indefinite Support Vector Machines
- Tangent Space Sensitivity and Distribution of Linear Regions in ReLU Networks
- Noise Stability Regularization for Improving BERT Fine-tuning
- A Theoretical Analysis of Learning with Noisily Labeled Data