Assessing Generalization of SGD via Disagreement
arXiv:2106.13799
Abstract
We empirically show that the test error of deep networks can be estimated by simply training the same architecture on the same training set but with a different run of Stochastic Gradient Descent (SGD), and measuring the disagreement rate between the two networks on unlabeled test data. This builds on -- and is a stronger version of -- the observation in Nakkiran & Bansal '20, which requires the second run to be on an altogether fresh training set. We further theoretically show that this peculiar phenomenon arises from the \emph{well-calibrated} nature of \emph{ensembles} of SGD-trained models. This finding not only provides a simple empirical measure to directly predict the test error using unlabeled test data, but also establishes a new conceptual connection between generalization and calibration.
References in corpus (13)
- On Calibration of Modern Neural Networks
- A Closer Look at Memorization in Deep Networks
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
- Evaluating model calibration in classification
- Generalization in Deep Networks: The Role of Distance from Initialization
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- Don't Just Blame Over-parametrization for Over-confidence: Theoretical Analysis of Calibration in Binary Classification
- Representation Based Complexity Measures for Predicting Generalization in Deep Learning
- Distributional Generalization: A New Kind of Generalization
- Moment Multicalibration for Uncertainty Estimation
- Should Ensemble Members Be Calibrated?
- Early-stopped neural networks are consistent
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
Cited by in corpus (7)
- Understanding Deep Learning via Decision Boundary
- A Framework for Cluster and Classifier Evaluation in the Absence of Reference Labels
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training Ensembles
- Evaluating Robustness to Dataset Shift via Parametric Robustness Sets
- On Predicting Generalization using GANs
- Explaining generalization in deep learning: progress and fundamental limits
- Toward Auto-evaluation with Confidence-based Category Relation-aware Regression