Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
arXiv:1708.03395 · doi:10.1073/pnas.1802705116
Abstract
Generalized linear models (GLMs) arise in high-dimensional machine learning, statistics, communications and signal processing. In this paper we analyze GLMs when the data matrix is random, as relevant in problems such as compressed sensing, error-correcting codes or benchmark models in neural networks. We evaluate the mutual information (or "free entropy") from which we deduce the Bayes-optimal estimation and generalization errors. Our analysis applies to the high-dimensional limit where both the number of samples and the dimension are large and their ratio is fixed. Non-rigorous predictions for the optimal errors existed for special cases of GLMs, e.g. for the perceptron, in the field of statistical physics based on the so-called replica method. Our present paper rigorously establishes those decades old conjectures and brings forward their algorithmic interpretation in terms of performance of the generalized approximate message-passing algorithm. Furthermore, we tightly characterize, for many learning problems, regions of parameters for which this algorithm achieves the optimal performance, and locate the associated sharp phase transitions separating learnable and non-learnable regions. We believe that this random version of GLMs can serve as a challenging benchmark for multi-purpose algorithms. This paper is divided in two parts that can be read independently: The first part (main part) presents the model and main results, discusses some applications and sketches the main ideas of the proof. The second part (supplementary informations) is much more detailed and provides more examples as well as all the proofs.
101 pages, 5 figures
References in corpus (9)
- Opening the Black Box of Deep Neural Networks via Information
- Compressive Phase Retrieval via Generalized Approximate Message Passing
- Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
- The Mutual Information in Random Linear Estimation
- Inference from correlated patterns: a unified theory for perceptron learning and linear vector channels
- Rethinking generalization requires revisiting old ideas: statistical mechanics approaches and complex learning behavior
- Structured signal recovery from quadratic measurements: Breaking sample complexity barriers via nonconvex optimization
- Toward Fast Reliable Communication at Rates Near Capacity with Gaussian Noise
- Phase transitions in spiked matrix estimation: information-theoretic analysis
Cited by in corpus (83)
- Machine learning and the physical sciences
- Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
- Entropy and mutual information in models of deep neural networks
- Modelling the influence of data structure on learning in neural networks: the hidden manifold model
- Generalisation error in learning with random features and the hidden manifold model
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- The distribution of the Lasso: Uniform control over sparse balls and adaptive parameter tuning
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- The adaptive interpolation method for proving replica formulas. Applications to the Curie-Weiss and Wigner spike models
- The Lasso with general Gaussian designs with applications to hypothesis testing
- High-temperature Expansions and Message Passing Algorithms
- Asymptotic learning curves of kernel methods: empirical data v.s. Teacher-Student paradigm
- Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
- Disordered Systems Insights on Computational Hardness
- Marvels and Pitfalls of the Langevin Algorithm in Noisy High-dimensional Inference
- Fundamental limits to learning closed-form mathematical models from data
- Annealing and replica-symmetry in Deep Boltzmann Machines
- The committee machine: Computational to statistical gaps in learning a two-layers neural network
- The Asymptotic Distribution of the MLE in High-dimensional Logistic Models: Arbitrary Covariance
- Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising
- Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
- Phase retrieval in high dimensions: Statistical and computational phase transitions
- Reducibility and Statistical-Computational Gaps from Secret Leakage
- High-dimensional rank-one nonsymmetric matrix decomposition: the spherical case
- Asymptotic errors for convex penalized linear regression beyond Gaussian matrices
- Learning curves for the multi-class teacher-student perceptron
- Fundamental Limits of Ridge-Regularized Empirical Risk Minimization in High Dimensions
- Fundamental problems in statistical physics XIV: Lecture on Machine Learning
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions
- All-or-nothing statistical and computational phase transitions in sparse spiked matrix estimation
- The estimation error of general first order methods
- Online stochastic gradient descent on non-convex losses from high-dimensional inference
- Deep learning via message passing algorithms based on belief propagation
- On the Universality of Noiseless Linear Estimation with Respect to the Measurement Matrix
- Macroscopic Analysis of Vector Approximate Message Passing in a Model Mismatch Setting
- Approximate Message Passing for orthogonally invariant ensembles: Multivariate non-linearities and spectral initialization
- Strong replica symmetry for high-dimensional disordered log-concave Gibbs measures
- Landscape Complexity for the Empirical Risk of Generalized Linear Models
- Thresholds of descending algorithms in inference problems
- Bayes-Optimal Estimation in Generalized Linear Models via Spatial Coupling
- On the interplay between data structure and loss function in classification problems
- Mutual information for low-rank even-order symmetric tensor estimation
- Concentration of multi-overlaps for random ferromagnetic spin models
- Theoretical characterization of uncertainty in high-dimensional linear classification
- Gardner formula for Ising perceptron models at small densities
- Prediction Errors for Penalized Regressions based on Generalized Approximate Message Passing
- Gibbs Sampling the Posterior of Neural Networks
- Bilinear Generalized Vector Approximate Message Passing
- Approximate Message Passing with Rigorous Guarantees for Pooled Data and Quantitative Group Testing
- Exact results on high-dimensional linear regression via statistical physics
- Mutual Information for the Stochastic Block Model by the Adaptive Interpolation Method
- Inference in Multi-Layer Networks with Matrix-Valued Unknowns
- Analysis of Bayesian Inference Algorithms by the Dynamical Functional Approach
- CAD: Debiasing the Lasso with inaccurate covariate model
- Neural-prior stochastic block model
- Minimum -norm interpolators: Precise asymptotics and multiple descent
- High-dimensional Asymptotics of VAEs: Threshold of Posterior Collapse and Dataset-Size Dependence of Rate-Distortion Curve
- Semi-analytic approximate stability selection for correlated data in generalized linear models
- Nonasymptotic Guarantees for Spiked Matrix Recovery with Generative Priors
- Generalization Error of Generalized Linear Models in High Dimensions
- Solvable Model for Inheriting the Regularization through Knowledge Distillation
- The Geometry of Over-parameterized Regression and Adversarial Perturbations
- Asymptotic mutual information in quadratic estimation problems over compact groups
- Analysis of Diffusion Models for Manifold Data
- Generalized Approximate Survey Propagation for High-Dimensional Estimation
- On the Cryptographic Hardness of Learning Single Periodic Neurons
- Free energy of multi-layer generalized linear models
- Blind calibration for compressed sensing: State evolution and an online algorithm
- Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
- High-dimensional inference: a statistical mechanics perspective
- Implicit Bias of Linear RNNs
- Construction of optimal spectral methods in phase retrieval
- Large dimensional analysis of general margin based classification methods
- Nonequilibrium thermodynamics of self-supervised learning
- Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation
- LASSO risk and phase transition under dependence
- Good Classifiers are Abundant in the Interpolating Regime
- Bias-variance decomposition of overparameterized regression with random linear features
- Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation
- Fundamental limits and algorithms for sparse linear regression with sublinear sparsity
- Analyzing Training Using Phase Transitions in Entropy---Part I: General Theory
- Statistical mechanics of extensive-width Bayesian neural networks near interpolation
- Inference in Spreading Processes with Neural-Network Priors