Classification vs regression in overparameterized regimes: Does the loss function matter?
arXiv:2005.08054
Abstract
We compare classification and regression tasks in an overparameterized linear model with Gaussian features. On the one hand, we show that with sufficient overparameterization all training points are support vectors: solutions obtained by least-squares minimum-norm interpolation, typically used for regression, are identical to those produced by the hard-margin support vector machine (SVM) that minimizes the hinge loss, typically used for training classifiers. On the other hand, we show that there exist regimes where these interpolating solutions generalize well when evaluated by the 0-1 test loss function, but do not generalize if evaluated by the square loss function, i.e. they approach the null risk. Our results demonstrate the very different roles and properties of loss functions used at the training phase (optimization) and the testing phase (generalization).
References in corpus (5)
- Understanding deep learning requires rethinking generalization
- More Data Can Hurt for Linear Regression: Sample-wise Double Descent
- Understanding overfitting peaks in generalization error: Analytical risk curves for and penalized interpolation
- Minimizing The Misclassification Error Rate Using a Surrogate Convex Loss
- Risk of the Least Squares Minimum Norm Estimator under the Spike Covariance Model
Cited by in corpus (30)
- Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime
- Interpolating Classifiers Make Few Mistakes
- Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
- Benign Overfitting of Constant-Stepsize SGD for Linear Regression
- Label-Imbalanced and Group-Sensitive Classification under Overparameterization
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian Mixtures
- Towards an Understanding of Benign Overfitting in Neural Networks
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- On the interplay between data structure and loss function in classification problems
- Towards Understanding the Data Dependency of Mixup-style Training
- Implicit Regularization via Neural Feature Alignment
- When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
- On the proliferation of support vectors in high dimensions
- Understanding and Mitigating Accuracy Disparity in Regression
- When does gradient descent with logistic loss find interpolating two-layer networks?
- Double Descent and Other Interpolation Phenomena in GANs
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- Multiplicative Reweighting for Robust Neural Network Optimization
- Prediction in latent factor regression: Adaptive PCR and beyond
- Minimax Supervised Clustering in the Anisotropic Gaussian Mixture Model: A new take on Robust Interpolation
- Explaining generalization in deep learning: progress and fundamental limits
- Interpolation can hurt robust generalization even when there is no noise
- Distilling Double Descent
- Classification and Adversarial examples in an Overparameterized Linear Model: A Signal Processing Perspective
- Memorization in Deep Neural Networks: Does the Loss Function matter?
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- The Interplay Between Implicit Bias and Benign Overfitting in Two-Layer Linear Networks
- How rotational invariance of common kernels prevents generalization in high dimensions
- On Generalization of Adaptive Methods for Over-parameterized Linear Regression
- On the Regularization of Autoencoders