Finite Versus Infinite Neural Networks: an Empirical Study
arXiv:2007.15801
Abstract
We perform a careful, thorough, and large scale empirical study of the correspondence between wide neural networks and kernel methods. By doing so, we resolve a variety of open questions related to the study of infinitely wide neural networks. Our experimental results include: kernel methods outperform fully-connected finite-width networks, but underperform convolutional finite width networks; neural network Gaussian process (NNGP) kernels frequently outperform neural tangent (NT) kernels; centered and ensembled finite networks have reduced posterior variance and behave more similarly to infinite networks; weight decay and the use of a large learning rate break the correspondence between finite and infinite networks; the NTK parameterization outperforms the standard parameterization for finite width networks; diagonal regularization of kernels acts similarly to early stopping; floating point precision limits kernel performance beyond a critical dataset size; regularized ZCA whitening improves accuracy; finite network performance depends non-monotonically on width in ways not captured by double descent phenomena; equivariance of CNNs is only beneficial for narrow networks far from the kernel regime. Our experiments additionally motivate an improved layer-wise scaling for weight decay which improves generalization in finite-width networks. Finally, we develop improved best practices for using NNGP and NT kernels for prediction, including a novel ensembling technique. Using these best practices we achieve state-of-the-art results on CIFAR-10 classification for kernels corresponding to each architecture class we consider.
17+11 pages; v2 references added, minor improvements
References in corpus (50)
- The NumPy array: a structure for efficient numerical computation
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Towards A Rigorous Science of Interpretable Machine Learning
- End to End Learning for Self-Driving Cars
- Popular Ensemble Methods: An Empirical Study
- Equality of Opportunity in Supervised Learning
- Reconciling modern machine learning practice and the bias-variance trade-off
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- An Analysis of Deep Neural Network Models for Practical Applications
- A Convergence Theory for Deep Learning via Over-Parameterization
- Machine Learning Methods for Attack Detection in the Smart Grid
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- On Lazy Training in Differentiable Programming
- Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
- Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
- Bayesian Deep Learning and a Probabilistic Perspective of Generalization
- Scaling description of generalization with number of parameters in deep learning
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
- An Empirical Model of Large-Batch Training
- Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures
- Enhanced Convolutional Neural Tangent Kernels
- SGD Learns the Conjugate Kernel Class of the Network
- The large learning rate phase of deep learning: the catapult mechanism
- Steps Toward Deep Kernel Methods from Infinite Neural Networks
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- On the Selection of Initialization and Activation Function for Deep Neural Networks
- A Fine-Grained Spectral Perspective on Neural Networks
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Neural Kernels Without Tangents
- Infinite attention: NNGP and NTK for deep attention networks
- On the training dynamics of deep networks with regularization
- Why do Larger Models Generalize Better? A Theoretical Perspective via the XOR Problem
- Deeper Connections between Neural Networks and Gaussian Processes Speed-up Active Learning
- Finite size corrections for neural network Gaussian processes
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
- Function Space Particle Optimization for Bayesian Neural Networks
- Asymptotics of Wide Convolutional Neural Networks
- The Effect of Network Width on Stochastic Gradient Descent and Generalization: an Empirical Study
- On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization
- On the infinite width limit of neural networks with a standard parameterization
- A Gaussian Process perspective on Convolutional Neural Networks
- On the expected behaviour of noise regularised deep neural networks as Gaussian processes
- Infinitely Wide Graph Convolutional Networks: Semi-supervised Learning via Gaussian Processes
- Critical initialisation for deep signal propagation in noisy rectifier neural networks
- A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off
- The role of a layer in deep neural networks: a Gaussian Process perspective
Cited by in corpus (10)
- A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit
- Wide and Deep Neural Networks Achieve Optimality for Classification
- Depth induces scale-averaging in overparameterized linear Bayesian neural networks
- Deep neural networks have an inbuilt Occam's razor
- Metric Flows with Neural Networks
- Infinite Neural Network Quantum States: Entanglement and Training Dynamics
- Relative stability toward diffeomorphisms indicates performance in deep nets
- A Dynamical View on Optimization Algorithms of Overparameterized Neural Networks
- Are wider nets better given the same number of parameters?
- A note on regularised NTK dynamics with an application to PAC-Bayesian training