A Random Matrix Analysis of Random Fourier Features: Beyond the Gaussian Kernel, a Precise Phase Transition, and the Corresponding Double Descent
arXiv:2006.05013 · doi:10.1088/1742-5468/ac3a77
Abstract
This article characterizes the exact asymptotics of random Fourier feature (RFF) regression, in the realistic setting where the number of data samples , their dimension , and the dimension of feature space are all large and comparable. In this regime, the random RFF Gram matrix no longer converges to the well-known limiting Gaussian kernel matrix (as it does when alone), but it still has a tractable behavior that is captured by our analysis. This analysis also provides accurate estimates of training and test regression errors for large . Based on these estimates, a precise characterization of two qualitatively different phases of learning, including the phase transition between them, is provided; and the corresponding double descent test error curve is derived from this phase transition behavior. These results do not depend on strong assumptions on the data distribution, and they perfectly match empirical results on real-world data sets.
NeurIPS 2020
References in corpus (9)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Reconciling modern machine learning practice and the bias-variance trade-off
- Benign Overfitting in Linear Regression
- On Lazy Training in Differentiable Programming
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Exact expressions for double descent and implicit regularization via surrogate random design
- On Random Matrices Arising in Deep Neural Networks. Gaussian Case
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networks
- Heavy-Tailed Universality Predicts Trends in Test Accuracies for Very Large Pre-Trained Deep Neural Networks
Cited by in corpus (7)
- On the interplay between data structure and loss function in classification problems
- Conditioning of Random Feature Matrices: Double Descent and Generalization Error
- Benign Overfitting and Noisy Features
- Learning from learning machines: a new generation of AI technology to meet the needs of science
- Asymptotics of Ridge Regression in Convolutional Models
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models
- Good Classifiers are Abundant in the Interpolating Regime