A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks
arXiv:1707.09564
Abstract
We present a generalization bound for feedforward neural networks in terms of the product of the spectral norm of the layers and the Frobenius norm of the weights. The generalization bound is derived using a PAC-Bayes analysis.
Accepted to ICLR 2018
Cited by in corpus (130)
- Neural Lander: Stable Drone Landing Control using Learned Dynamics
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- What is being transferred in transfer learning?
- Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
- Are Anchor Points Really Indispensable in Label-Noise Learning?
- Bayesian Deep Learning and a Probabilistic Perspective of Generalization
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- Contraction Theory for Nonlinear Stability Analysis and Learning-based Control: A Tutorial Overview
- When Do Extended Physics-Informed Neural Networks (XPINNs) Improve Generalization?
- The Modern Mathematics of Deep Learning
- A Theoretical Analysis of Deep Q-Learning
- Predicting the Generalization Gap in Deep Networks with Margin Distributions
- Rademacher Complexity for Adversarially Robust Generalization
- A Priori Estimates of the Population Risk for Two-layer Neural Networks
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Are All Layers Created Equal?
- Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- The Pitfalls of Simplicity Bias in Neural Networks
- Sorting out Lipschitz function approximation
- User-friendly introduction to PAC-Bayes bounds
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- A Constructive Prediction of the Generalization Error Across Scales
- An analytic theory of generalization dynamics and transfer learning in deep linear networks
- Identity Crisis: Memorization and Generalization under Extreme Overparameterization
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Identifying Generalization Properties in Neural Networks
- Tighter risk certificates for neural networks
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex Learning
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
- Why Spectral Normalization Stabilizes GANs: Analysis and Improvements
- Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
- Subgroup Generalization and Fairness of Graph Neural Networks
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- The Implicit and Explicit Regularization Effects of Dropout
- Theoretical properties of the global optimizer of two layer neural network
- Understanding Generalization through Visualizations
- Assessing Generalization of SGD via Disagreement
- On the role of data in PAC-Bayes bounds
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path
- Algorithmic Regularization in Over-parameterized Matrix Sensing and Neural Networks with Quadratic Activations
- On the distance between two neural networks and the stability of learning
- The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
- Inductive Biases and Variable Creation in Self-Attention Mechanisms
- SiPPing Neural Networks: Sensitivity-informed Provable Pruning of Neural Networks
- PAC-Bayes with Backprop
- Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
- Generalization Error Bounds with Probabilistic Guarantee for SGD in Nonconvex Optimization
- LocalDrop: A Hybrid Regularization for Deep Neural Networks
- Minnorm training: an algorithm for training over-parameterized deep neural networks
- Generalization Bounds for Neural Belief Propagation Decoders
- Convergence Rates of Variational Inference in Sparse Deep Learning
- Dropout: Explicit Forms and Capacity Control
- Depth separation for reduced deep networks in nonlinear model reduction: Distilling shock waves in nonlinear hyperbolic problems
- Understanding Generalization in Deep Learning via Tensor Methods
- Hypothesis Set Stability and Generalization
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural network
- Generalization Guarantees for Imitation Learning
- How Many Samples are Needed to Estimate a Convolutional or Recurrent Neural Network?
- Data-dependent PAC-Bayes priors via differential privacy
- Learning Curves for Analysis of Deep Networks
- Information-Theoretic Generalization Bounds for Stochastic Gradient Descent
- Is SGD a Bayesian sampler? Well, almost
- Taxonomizing local versus global structure in neural network loss landscapes
- Gradient-Free Learning Based on the Kernel and the Range Space
- Stationary Points of Shallow Neural Networks with Quadratic Activation Function
- SALR: Sharpness-aware Learning Rate Scheduler for Improved Generalization
- PAC-Bayesian Margin Bounds for Convolutional Neural Networks
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Evaluation of Complexity Measures for Deep Learning Generalization in Medical Image Analysis
- Analytic Network Learning
- Train simultaneously, generalize better: Stability of gradient-based minimax learners
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- Distance-Based Regularisation of Deep Networks for Fine-Tuning
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- PAC-Bayes Analysis Beyond the Usual Bounds
- Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds
- Sample Complexity Bounds for Recurrent Neural Networks with Application to Combinatorial Graph Problems
- PAC-Bayesian Contrastive Unsupervised Representation Learning
- Generalization Guarantees for Neural Architecture Search with Train-Validation Split
- A Deep Conditioning Treatment of Neural Networks
- A Dynamical View on Optimization Algorithms of Overparameterized Neural Networks
- Information-Theoretic Local Minima Characterization and Regularization
- Misclassification bounds for PAC-Bayesian sparse deep learning
- Dichotomize and Generalize: PAC-Bayesian Binary Activated Deep Neural Networks
- What Information Does a ResNet Compress?
- Efron-Stein PAC-Bayesian Inequalities
- A Limitation of the PAC-Bayes Framework
- Norm-based generalisation bounds for multi-class convolutional neural networks
- De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors
- On a Sparse Shortcut Topology of Artificial Neural Networks
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural Networks
- Why flatness does and does not correlate with generalization for deep neural networks
- Self-Regularity of Non-Negative Output Weights for Overparameterized Two-Layer Neural Networks
- Abstraction Mechanisms Predict Generalization in Deep Neural Networks
- Improve Generalization and Robustness of Neural Networks via Weight Scale Shifting Invariant Regularizations
- Learning with Noisy Labels by Efficient Transition Matrix Estimation to Combat Label Miscorrection
- PAC-Bayesian Transportation Bound
- Initialization and Regularization of Factorized Neural Layers
- Stable Rank Normalization for Improved Generalization in Neural Networks and GANs
- The Traveling Observer Model: Multi-task Learning Through Spatial Variable Embeddings
- Structured Dropout Variational Inference for Bayesian Neural Networks
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and Implications
- Understanding Deep Architectures with Reasoning Layer
- On the Robustness and Generalization of Deep Learning Driven Full Waveform Inversion
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- Why Do Better Loss Functions Lead to Less Transferable Features?
- Revisiting minimum description length complexity in overparameterized models
- Harmonic Decompositions of Convolutional Networks
- Provably Training Overparameterized Neural Network Classifiers with Non-convex Constraints
- Leveraging Local Variation in Data: Sampling and Weighting Schemes for Supervised Deep Learning
- The role of invariance in spectral complexity-based generalization bounds
- Uniform Convergence, Adversarial Spheres and a Simple Remedy
- Limiting Network Size within Finite Bounds for Optimization
- Generalization Performance of Empirical Risk Minimization on Over-parameterized Deep ReLU Nets
- Rethinking Breiman's Dilemma in Neural Networks: Phase Transitions of Margin Dynamics
- Generalisation under gradient descent via deterministic PAC-Bayes
- Comparing Comparators in Generalization Bounds
- On the Structural Sensitivity of Deep Convolutional Networks to the Directions of Fourier Basis Functions
- PAC-Bayes-Chernoff bounds for unbounded losses
- Gentle Local Robustness implies Generalization
- FeTa: A DCA Pruning Algorithm with Generalization Error Guarantees
- MSR-DARTS: Minimum Stable Rank of Differentiable Architecture Search
- Towards Understanding Generalization via Decomposing Excess Risk Dynamics
- Recent Advances in Large Margin Learning
- Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform Stability
- Information Theoretic Lower Bounds for Feed-Forward Fully-Connected Deep Networks
- A note on regularised NTK dynamics with an application to PAC-Bayesian training