Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
arXiv:1910.04284
Abstract
For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization. Unfortunately, for deep models, this relationship is less clear: existing analyses of the output margin give complicated bounds which sometimes depend exponentially on depth. In this work, we propose to instead analyze a new notion of margin, which we call the "all-layer margin." Our analysis reveals that the all-layer margin has a clear and direct relationship with generalization for deep models. This enables the following concrete applications of the all-layer margin: 1) by analyzing the all-layer margin, we obtain tighter generalization bounds for neural nets which depend on Jacobian and hidden layer norms and remove the exponential dependency on depth 2) our neural net results easily translate to the adversarially robust setting, giving the first direct analysis of robust test error for deep networks, and 3) we present a theoretically inspired training algorithm for increasing the all-layer margin. Our algorithm improves both clean and adversarially robust test performance over strong baselines in practice.
Code for all-layer margin optimization is available at the following link: https://github.com/cwein3/all-layer-margin-opt. Version 4: Re-organized proofs for more clarity
References in corpus (25)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Learning Face Representation by Joint Identification-Verification
- Understanding deep learning requires rethinking generalization
- Theoretically Principled Trade-off between Robustness and Accuracy
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Benign Overfitting in Linear Regression
- Adversarially Robust Generalization Requires More Data
- Robustness May Be at Odds with Accuracy
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
- Adversarial Training Can Hurt Generalization
- Predicting the Generalization Gap in Deep Networks with Margin Distributions
- Rademacher Complexity for Adversarially Robust Generalization
- To understand deep learning we need to understand kernel learning
- Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
- Stronger generalization bounds for deep nets via a compression approach
- Norm-Based Capacity Control in Neural Networks
- Generalizable Adversarial Training via Spectral Normalization
- Risk and parameter convergence of logistic regression
- Adversarial Examples Improve Image Recognition
- VC Classes are Adversarially Robustly Learnable, but Only Improperly
- Adversarial Risk Bounds via Function Transformation
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Adversarial Margin Maximization Networks
- Implicit Rugosity Regularization via Data Augmentation
Cited by in corpus (20)
- MOPO: Model-based Offline Policy Optimization
- Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
- Fantastic Generalization Measures and Where to Find Them
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
- Optimal Regularization Can Mitigate Double Descent
- The Curious Case of Adversarially Robust Models: More Data Can Help, Double Descend, or Hurt Generalization
- The Implicit and Explicit Regularization Effects of Dropout
- Heteroskedastic and Imbalanced Deep Learning with Adaptive Regularization
- Generalization bounds for deep learning
- Angular Visual Hardness
- Adversarial Feature Augmentation and Normalization for Visual Recognition
- A Theory of Label Propagation for Subpopulation Shift
- Data-Efficient GAN Training Beyond (Just) Augmentations: A Lottery Ticket Perspective
- Label Noise SGD Provably Prefers Flat Global Minimizers
- Train simultaneously, generalize better: Stability of gradient-based minimax learners
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- Minimum sharpness: Scale-invariant parameter-robustness of neural networks
- Recent Advances in Large Margin Learning
- Towards Understanding Generalization via Decomposing Excess Risk Dynamics
- Noisy Feature Mixup