Relative Flatness and Generalization
arXiv:2001.00939
Abstract
Flatness of the loss curve is conjectured to be connected to the generalization ability of machine learning models, in particular neural networks. While it has been empirically observed that flatness measures consistently correlate strongly with generalization, it is still an open theoretical problem why and under which circumstances flatness is connected to generalization, in particular in light of reparameterizations that change certain flatness measures but leave generalization unchanged. We investigate the connection between flatness and generalization by relating it to the interpolation from representative data, deriving notions of representativeness, and feature robustness. The notions allow us to rigorously connect flatness and generalization and to identify conditions under which the connection holds. Moreover, they give rise to a novel, but natural relative flatness measure that correlates strongly with generalization, simplifies to ridge regression for ordinary least squares, and solves the reparameterization issue.
The first two authors made equal contribution; Accepted for publication at NeurIPS 2021; arXiv admin note: substantial text overlap with arXiv:1912.00058
References in corpus (9)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Reconciling modern machine learning practice and the bias-variance trade-off
- Benign Overfitting in Linear Regression
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Uniform convergence may be unable to explain generalization in deep learning
- Norm-Based Capacity Control in Neural Networks
- Theory of Deep Learning IIb: Optimization Properties of SGD
- A Scale Invariant Flatness Measure for Deep Network Minima
- Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter Corruption