A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
arXiv:1912.00827
Abstract
One of the distinguishing characteristics of modern deep learning systems is that they typically employ neural network architectures that utilize enormous numbers of parameters, often in the millions and sometimes even in the billions. While this paradigm has inspired significant research on the properties of large networks, relatively little work has been devoted to the fact that these networks are often used to model large complex datasets, which may themselves contain millions or even billions of constraints. In this work, we focus on this high-dimensional regime in which both the dataset size and the number of features tend to infinity. We analyze the performance of random feature regression with features for a random weight matrix and random bias vector , obtaining exact formulae for the asymptotic training and test errors for data generated by a linear teacher model. The role of the bias can be understood as parameterizing a distribution over activation functions, and our analysis directly generalizes to such distributions, even those not expressible with a traditional additive bias. Intriguingly, we find that a mixture of nonlinearities can improve both the training and test errors over the best single nonlinearity, suggesting that mixtures of nonlinearities might be useful for approximate kernel methods or neural network architecture design.
References in corpus (5)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
- Kernel Alignment Risk Estimator: Risk Prediction from Training Data
- Understanding Double Descent Requires a Fine-Grained Bias-Variance Decomposition
- The universality principle for spectral distributions of sample covariance matrices
Cited by in corpus (10)
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- What causes the test error? Going beyond bias-variance via ANOVA
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networks
- Analysis of One-Hidden-Layer Neural Networks via the Resolvent Method
- Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature model
- Mixed Moments for the Product of Ginibre Matrices
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU Networks
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Covariate Shift in High-Dimensional Random Feature Regression
- Large-Dimensional Random Matrix Theory and Its Applications in Deep Learning and Wireless Communications