Are Flat Minima an Illusion?
arXiv:2605.05209
Abstract
Flat minima are an account of why deep networks generalise. However flatness is a matter of form (parameters), while generalisation is of function. The same function can be a result of many different parameterisations. I demonstrate this by rescaling ReLU networks, changing raw Hessian trace by up to times while every prediction remains fixed. Raw curvature cannot identify a function-level explanation. Previous theoretical work traced generalisation to the weakness of constraints implied by function, meaning the freedom a model retains within the bounds of what it has learned to be correct. A policy is weaker when more future commitments remain compatible with what it has learned, allowing more freedom to adapt. To measure this for neural networks, I freeze the last hidden representation and ask whether each of 512 sampled label bundles can be met by a replacement affine classifier. The resulting joint completion score is invariant under invertible linear mixing and translation of feature coordinates. Across two predeclared cohorts of 100 networks, it predicts held-out accuracy with rank correlations and . Raw Hessian trace and relative flatness have no multiplicity-corrected association. To put it provocatively, freedom is correlated with adaptability, while flatness is a matter of description.
27 pages, 1 figure. Major revision adds an affine-invariant joint completion score, PAC-Bayes certificates, a task-alignment theorem, three predeclared 100-network cohorts, a random-label control, and expanded references. Submitted to JMLR