Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
arXiv:1602.05897
Abstract
We develop a general duality between neural networks and compositional kernels, striving towards a better understanding of deep learning. We show that initial representations generated by common random initializations are sufficiently rich to express all functions in the dual kernel space. Hence, though the training objective is hard to optimize in the worst case, the initial weights form a good starting point for optimization. Our dual view also reveals a pragmatic and aesthetic perspective of neural networks and underscores their expressive power.
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- The Loss Surfaces of Multilayer Networks
- Convolutional Kernel Networks
- Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Norm-Based Capacity Control in Neural Networks
- On the Computational Efficiency of Training Neural Networks
- Steps Toward Deep Kernel Methods from Infinite Neural Networks
- Deep Convolutional Networks are Hierarchical Kernel Machines
- On the Equivalence between Kernel Quadrature Rules and Random Feature Expansions
Cited by in corpus (122)
- Deep Learning with Differential Privacy
- Deep Neural Networks as Gaussian Processes
- Deep Reinforcement Learning: An Overview
- On Exact Computation with an Infinitely Wide Neural Net
- Convergence Analysis of Two-layer Neural Networks with ReLU Activation
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
- Learning One-hidden-layer Neural Networks with Landscape Design
- On the Convergence Rate of Training Recurrent Neural Networks
- AdaNet: Adaptive Structural Learning of Artificial Neural Networks
- What Can ResNet Learn Efficiently, Going Beyond Kernels?
- Denoising and Regularization via Exploiting the Structural Bias of Convolutional Generators
- On Characterizing the Capacity of Neural Networks using Algebraic Topology
- Neural Tangents: Fast and Easy Infinite Neural Networks in Python
- Kernel and Rich Regimes in Overparametrized Models
- Convergence of Adversarial Training in Overparametrized Neural Networks
- A Mean Field Theory of Batch Normalization
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- Asymptotics of Wide Networks from Feynman Diagrams
- Spurious Valleys in Two-layer Neural Network Optimization Landscapes
- Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes
- Finite Versus Infinite Neural Networks: an Empirical Study
- A Fine-Grained Spectral Perspective on Neural Networks
- Neural Kernels Without Tangents
- Mean Field Residual Networks: On the Edge of Chaos
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? -- A Neural Tangent Kernel Perspective
- Mathematical Models of Overparameterized Neural Networks
- On the Power and Limitations of Random Features for Understanding Neural Networks
- Simple, Fast, and Flexible Framework for Matrix Completion with Infinite Width Neural Networks
- Deriving Neural Architectures from Sequence and Graph Kernels
- Universality Laws for High-Dimensional Learning with Random Features
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Tensor Programs III: Neural Matrix Laws
- Neural Networks Learning and Memorization with (almost) no Over-Parameterization
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networks
- Towards NNGP-guided Neural Architecture Search
- Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics
- Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural Networks
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks
- A Generalizable and Accessible Approach to Machine Learning with Global Satellite Imagery
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- SGD Learns One-Layer Networks in WGANs
- A Correspondence Between Random Neural Networks and Statistical Field Theory
- Decoupling Gating from Linearity
- Zen-NAS: A Zero-Shot NAS for High-Performance Deep Image Recognition
- On the Similarity between the Laplace and Neural Tangent Kernels
- Effect of Activation Functions on the Training of Overparametrized Neural Nets
- Learning Parities with Neural Networks
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- Order and Chaos: NTK views on DNN Normalization, Checkerboard and Boundary Artifacts
- On the expected behaviour of noise regularised deep neural networks as Gaussian processes
- Algorithms and SQ Lower Bounds for PAC Learning One-Hidden-Layer ReLU Networks
- Deep Function Machines: Generalized Neural Networks for Topological Layer Expression
- Pathological spectra of the Fisher information metric and its variants in deep neural networks
- Invariance of Weight Distributions in Rectified MLPs
- Deep Equals Shallow for ReLU Networks in Kernel Regimes
- Variational Implicit Processes
- The Normalization Method for Alleviating Pathological Sharpness in Wide Neural Networks
- Early Stopping in Deep Networks: Double Descent and How to Eliminate it
- Generalized Leverage Score Sampling for Neural Networks
- Eigenvalue Decay Implies Polynomial-Time Learnability for Neural Networks
- On the Estimation of Derivatives Using Plug-in Kernel Ridge Regression Estimators
- Hardness of Learning Neural Networks with Natural Weights
- Random Features for Compositional Kernels
- Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks
- Symmetry & critical points for a model shallow neural network
- Richer priors for infinitely wide multi-layer perceptrons
- Wide Neural Networks with Bottlenecks are Deep Gaussian Processes
- Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
- Approximation and Learning with Deep Convolutional Models: a Kernel Perspective
- Deep Online Convex Optimization with Gated Games
- A Spectral Analysis of Dot-product Kernels
- Non-asymptotic approximations of neural networks by Gaussian processes
- Quantifying the Benefit of Using Differentiable Learning over Tangent Kernels
- Infinitely Wide Tensor Networks as Gaussian Process
- A Deep Conditioning Treatment of Neural Networks
- Wider Networks Learn Better Features
- Learning with convolution and pooling operations in kernel methods
- Forward Super-Resolution: How Can GANs Learn Hierarchical Generative Models for Real-World Distributions
- Which Minimizer Does My Neural Network Converge To?
- Building Bayesian Neural Networks with Blocks: On Structure, Interpretability and Uncertainty
- Kernel-based Translations of Convolutional Networks
- For Manifold Learning, Deep Neural Networks can be Locality Sensitive Hash Functions
- Fractional moment-preserving initialization schemes for training deep neural networks
- Scaling Neural Tangent Kernels via Sketching and Random Features
- Trap of Feature Diversity in the Learning of MLPs
- Generalization Error of Generalized Linear Models in High Dimensions
- Fundamental Tradeoffs in Distributionally Adversarial Training
- Randomness in Deconvolutional Networks for Visual Representation
- DNN-Based Topology Optimisation: Spatial Invariance and Neural Tangent Kernel
- Mean field theory for deep dropout networks: digging up gradient backpropagation deeply
- Conditional Deep Gaussian Processes: multi-fidelity kernel learning
- Making Method of Moments Great Again? -- How can GANs learn distributions
- Implicit Bias of Linear RNNs
- MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
- A Temporal Kernel Approach for Deep Learning with Continuous-time Information
- On the validity of kernel approximations for orthogonally-initialized neural networks
- Measure, Manifold, Learning, and Optimization: A Theory Of Neural Networks
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Properties of the After Kernel
- Learning curves for Gaussian process regression with power-law priors and targets
- Critical Percolation as a Framework to Analyze the Training of Deep Networks
- On the Provable Generalization of Recurrent Neural Networks
- Towards Understanding Learning in Neural Networks with Linear Teachers
- Untrained Graph Neural Networks for Denoising
- A Johnson--Lindenstrauss Framework for Randomly Initialized CNNs
- Uniform Generalization Bounds for Overparameterized Neural Networks
- Neural Optimization Kernel: Towards Robust Deep Learning
- Over-parametrized neural networks as under-determined linear systems
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- The Expressivity and Training of Deep Neural Networks: toward the Edge of Chaos?
- How rotational invariance of common kernels prevents generalization in high dimensions
- Reservoir Transformers
- Implicit Acceleration and Feature Learning in Infinitely Wide Neural Networks with Bottlenecks
- Random Features for the Neural Tangent Kernel
- From deep to Shallow: Equivalent Forms of Deep Networks in Reproducing Kernel Krein Space and Indefinite Support Vector Machines
- Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping