Steps Toward Deep Kernel Methods from Infinite Neural Networks
arXiv:1508.05133
Abstract
Contemporary deep neural networks exhibit impressive results on practical problems. These networks generalize well although their inherent capacity may extend significantly beyond the number of training examples. We analyze this behavior in the context of deep, infinite neural networks. We show that deep infinite layers are naturally aligned with Gaussian processes and kernel methods, and devise stochastic kernels that encode the information of these networks. We show that stability results apply despite the size, offering an explanation for their empirical success.
References in corpus (1)
Cited by in corpus (36)
- Deep Neural Networks as Gaussian Processes
- On Exact Computation with an Infinitely Wide Neural Net
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Extrapolating quantum observables with machine learning: Inferring multiple phase transitions from properties of a single phase
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes
- Finite Versus Infinite Neural Networks: an Empirical Study
- A Fine-Grained Spectral Perspective on Neural Networks
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- Recent advances in deep learning theory
- Deriving Neural Architectures from Sequence and Graph Kernels
- Tensor Programs III: Neural Matrix Laws
- Towards NNGP-guided Neural Architecture Search
- Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks
- A Gaussian Process perspective on Convolutional Neural Networks
- Approximate Inference Turns Deep Networks into Gaussian Processes
- On the expected behaviour of noise regularised deep neural networks as Gaussian processes
- Variational Implicit Processes
- Nonparametric Neural Networks
- Deep Function Machines: Generalized Neural Networks for Topological Layer Expression
- Batch Normalization Orthogonalizes Representations in Deep Random Networks
- Benign Overfitting and Noisy Features
- Wide Neural Networks with Bottlenecks are Deep Gaussian Processes
- Non-asymptotic approximations of neural networks by Gaussian processes
- Correlated Weights in Infinite Limits of Deep Convolutional Neural Networks
- Bayesian Learning of LF-MMI Trained Time Delay Neural Networks for Speech Recognition
- On Universal Approximation by Neural Networks with Uniform Guarantees on Approximation of Infinite Dimensional Maps
- Infinitely Wide Tensor Networks as Gaussian Process
- Label-Aware Neural Tangent Kernel: Toward Better Generalization and Local Elasticity
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations
- Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
- Implicit Acceleration and Feature Learning in Infinitely Wide Neural Networks with Bottlenecks
- On the relationship between multitask neural networks and multitask Gaussian Processes