Generalisation error in learning with random features and the hidden manifold model
arXiv:2002.09339 · doi:10.1088/1742-5468/ac3ae6
Abstract
We study generalised linear regression and classification for a synthetically generated dataset encompassing different problems of interest, such as learning with random features, neural networks in the lazy training regime, and the hidden manifold model. We consider the high-dimensional regime and using the replica method from statistical physics, we provide a closed-form expression for the asymptotic generalisation performance in these problems, valid in both the under- and over-parametrised regimes and for a broad choice of generalised linear model loss functions. In particular, we show how to obtain analytically the so-called double descent behaviour for logistic regression with a peak at the interpolation threshold, we illustrate the superiority of orthogonal against random Gaussian projections in learning with random features, and discuss the role played by correlations in the data generated by the hidden manifold model. Beyond the interest in these particular problems, the theoretical formalism introduced in this manuscript provides a path to further extensions to more complex tasks.
v2: ICML 2020 camera-ready
References in corpus (6)
- Exploring Generalization in Deep Learning
- Mean-field message-passing equations in the Hopfield model and its generalizations
- The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime
- The Gaussian equivalence of generative models for learning with shallow neural networks
- Kernel computations from large-scale random features obtained by Optical Processing Units
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian Mixtures
Cited by in corpus (31)
- Learning curves of generic features maps for realistic datasets with a teacher-student model
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Mapping of attention mechanisms to a generalized Potts model
- Asymptotic Errors for Teacher-Student Convex Generalized Linear Models (or : How to Prove Kabashima's Replica Formula)
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Learning through atypical "phase transitions" in overparameterized neural networks
- A Precise Performance Analysis of Learning with Random Features
- Universality Laws for High-Dimensional Learning with Random Features
- Probing transfer learning with a model of synthetic correlated datasets
- Gaussian Universality of Perceptrons with Random Labels
- On Data-Augmentation and Consistency-Based Semi-Supervised Learning
- On the Inherent Regularization Effects of Noise Injection During Training
- Distributional Generalization: A New Kind of Generalization
- Optimal regularizations for data generation with probabilistic graphical models
- When Does Preconditioning Help or Hurt Generalization?
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime
- Random features and polynomial rules
- Asymptotic theory of in-context learning by linear attention
- On the interplay between data structure and loss function in classification problems
- Theoretical characterization of uncertainty in high-dimensional linear classification
- Breaking the waves: asymmetric random periodic features for low-bitrate kernel machines
- Optimal Learning with Excitatory and Inhibitory synapses
- High-dimensional Asymptotics of VAEs: Threshold of Posterior Collapse and Dataset-Size Dependence of Rate-Distortion Curve
- Testing the Isotropic Cauchy Hypothesis
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data
- How isotropic kernels perform on simple invariants
- Statistical mechanics of extensive-width Bayesian neural networks near interpolation
- Ising Model Selection Using -Regularized Linear Regression: A Statistical Mechanics Analysis
- Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation
- Minimum complexity interpolation in random features models
- Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation