The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime
arXiv:1911.01544
Abstract
Modern machine learning classifiers often exhibit vanishing classification error on the training set. They achieve this by learning nonlinear representations of the inputs that maps the data into linearly separable classes. Motivated by these phenomena, we revisit high-dimensional maximum margin classification for linearly separable data. We consider a stylized setting in which data , are i.i.d. with a -dimensional Gaussian feature vector, and a label whose distribution depends on a linear combination of the covariates . While the Gaussian model might appear extremely simplistic, universality arguments can be used to show that the results derived in this setting also apply to the output of certain nonlinear featurization maps. We consider the proportional asymptotics with , and derive exact expressions for the limiting generalization error. We use this theory to derive two results of independent interest: Sufficient conditions on for `benign overfitting' that parallel previously derived conditions in the case of linear regression; An asymptotically exact expression for the generalization error when max-margin classification is used in conjunction with feature vectors produced by random one-layer neural networks.
97 pages; 12 pdf figures (A major revision; changed the title and included a new result on benign overfitting)
References in corpus (2)
Cited by in corpus (48)
- Generalisation error in learning with random features and the hidden manifold model
- Learning curves of generic features maps for realistic datasets with a teacher-student model
- The Gaussian equivalence of generative models for learning with shallow neural networks
- Classification vs regression in overparameterized regimes: Does the loss function matter?
- The Lasso with general Gaussian designs with applications to hypothesis testing
- Universality Laws for High-Dimensional Learning with Random Features
- A Precise Performance Analysis of Learning with Random Features
- Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- A Model of Double Descent for High-dimensional Binary Linear Classification
- Gaussian Universality of Perceptrons with Random Labels
- Fundamental Limits of Ridge-Regularized Empirical Risk Minimization in High Dimensions
- On the Optimal Weighted Regularization in Overparameterized Linear Regression
- Interpolating Classifiers Make Few Mistakes
- On the Inherent Regularization Effects of Noise Injection During Training
- Mean-Field Neural ODEs via Relaxed Optimal Control
- Theoretical Insights Into Multiclass Classification: A High-dimensional Asymptotic View
- Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
- Precise Tradeoffs in Adversarial Training for Linear Regression
- On the Generalization Effects of Linear Transformations in Data Augmentation
- When Does Preconditioning Help or Hurt Generalization?
- Label-Imbalanced and Group-Sensitive Classification under Overparameterization
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- On the interplay between data structure and loss function in classification problems
- On Uniform Convergence and Low-Norm Interpolation Learning
- Provable More Data Hurt in High Dimensional Least Squares Estimator
- When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
- Sharp Asymptotics and Optimal Performance for Inference in Binary Models
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Inference in Multi-Layer Networks with Matrix-Valued Unknowns
- Exact Gap between Generalization Error and Uniform Convergence in Random Feature Models
- When does gradient descent with logistic loss find interpolating two-layer networks?
- CAD: Debiasing the Lasso with inaccurate covariate model
- Bridging the Gap Between Adversarial Robustness and Optimization Bias
- Asymptotic Behavior of Adversarial Training in Binary Classification
- The Performance Analysis of Generalized Margin Maximizer (GMM) on Separable Data
- Generalization Error of Generalized Linear Models in High Dimensions
- Prediction in latent factor regression: Adaptive PCR and beyond
- Explaining generalization in deep learning: progress and fundamental limits
- Generalization error of minimum weighted norm and kernel interpolation
- Implicit Bias of Linear RNNs
- Asymptotic Risk of Overparameterized Likelihood Models: Double Descent Theory for Deep Neural Networks
- Training Efficiency and Robustness in Deep Learning
- Improved Complexities for Stochastic Conditional Gradient Methods under Interpolation-like Conditions
- Robustifying Binary Classification to Adversarial Perturbation
- Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective
- The Interplay Between Implicit Bias and Benign Overfitting in Two-Layer Linear Networks
- How isotropic kernels perform on simple invariants