Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation
arXiv:2510.24616 · doi:10.1103/56sb-pdh6
Abstract
For four decades statistical physics has been providing a framework to analyse neural networks. A long-standing question remained on its capacity to tackle deep learning models capturing rich feature learning effects, thus going beyond the narrow networks or kernel methods analysed until now. We positively answer through the study of the supervised learning of a multi-layer perceptron. Importantly, (i) its width scales as the input dimension, making it more prone to feature learning than ultra wide networks, and more expressive than narrow ones or ones with fixed embedding layers; and (ii) we focus on the challenging interpolation regime where the number of trainable parameters and data are comparable, which forces the model to adapt to the task. We consider the matched teacher-student setting. Therefore, we provide the fundamental limits of learning random deep neural network targets and identify the sufficient statistics describing what is learnt by an optimally trained network as the data budget increases. A rich phenomenology emerges with various learning transitions. With enough data, optimal performance is attained through the model's "specialisation" towards the target, but it can be hard to reach for training algorithms which get attracted by sub-optimal solutions predicted by the theory. Specialisation occurs inhomogeneously across layers, propagating from shallow towards deep ones, but also across neurons in each layer. Furthermore, deeper targets are harder to learn. Despite its simplicity, the Bayes-optimal setting provides insights on how the depth, non-linearity and finite (proportional) width influence neural networks in the feature learning regime that are potentially relevant in much more general settings.
33 pages, 20 figures + appendix. This submission supersedes both arXiv:2505.24849 and arXiv:2501.18530. v5 fixes minor typos
References in corpus (39)
- A Mean Field View of the Landscape of Two-Layers Neural Networks
- Statistical physics of inference: Thresholds and algorithms
- Bilinear Generalized Approximate Message Passing
- Random matrices: Universality of local eigenvalue statistics up to the edge
- Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
- Classification and Geometry of General Perceptual Manifolds
- Statistical Mechanics of Support Vector Networks
- Modelling the influence of data structure on learning in neural networks: the hidden manifold model
- Phase transitions and sample complexity in Bayes-optimal matrix factorization
- Spectral Bias and Task-Model Alignment Explain Generalization in Kernel Regression and Infinitely Wide Neural Networks
- Analysis of CDMA systems that are characterized by eigenvalue spectrum
- Generalisation error in learning with random features and the hidden manifold model
- Rotational invariant estimator for general noisy matrices
- A Theory of Solving TAP Equations for Ising Models with General Invariant Random Matrices
- Properties of the geometry of solutions and capacity of multi-layer neural networks with Rectified Linear Units activations
- Trainability and Accuracy of Neural Networks: An Interacting Particle System Approach
- Statistical Mechanics of Deep Linear Neural Networks: The Back-Propagating Kernel Renormalization
- Inference from correlated patterns: a unified theory for perceptron learning and linear vector channels
- High-temperature Expansions and Message Passing Algorithms
- A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit
- Notes on Matrix Models
- Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
- Instanton Approach to Large Harish-Chandra-Itzykson-Zuber Integrals
- Asymptotic Errors for Teacher-Student Convex Generalized Linear Models (or : How to Prove Kabashima's Replica Formula)
- Statistical limits of dictionary learning: random matrix theory and the spectral replica method
- Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising
- Statistical learning theory of structured data
- Statistical Mechanics of Dictionary Learning
- How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
- Beyond the storage capacity: data driven satisfiability transition
- On the large N limit of matrix integrals over the orthogonal group
- From complex to simple : hierarchical free-energy landscape renormalized in deep neural networks
- Macroscopic Analysis of Vector Approximate Message Passing in a Model Mismatch Setting
- Matrix denoising: Bayes-optimal estimators via low-degree polynomials
- Random features and polynomial rules
- Coding schemes in neural networks learning classification tasks
- Spatially heterogeneous learning by a deep student machine
- Correlations between hidden units in multilayer neural networks and replica symmetry breaking
- Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens