A Mean Field View of the Landscape of Two-Layers Neural Networks
arXiv:1804.06561 · doi:10.1073/pnas.1806579115
Abstract
Multi-layer neural networks are among the most powerful models in machine learning, yet the fundamental reasons for this success defy mathematical understanding. Learning a neural network requires to optimize a non-convex high-dimensional objective (risk function), a problem which is usually attacked using stochastic gradient descent (SGD). Does SGD converge to a global optimum of the risk or only to a local optimum? In the first case, does this happen because local minima are absent, or because SGD somehow avoids them? In the second, why do local minima reached by SGD have good generalization properties? In this paper we consider a simple case, namely two-layers neural networks, and prove that -in a suitable scaling limit- SGD dynamics is captured by a certain non-linear partial differential equation (PDE) that we call distributional dynamics (DD). We then consider several specific examples, and show how DD can be used to prove convergence of SGD to networks with nearly ideal generalization error. This description allows to 'average-out' some of the complexities of the landscape of neural networks, and can be used to prove a general convergence result for noisy SGD.
103 pages
References in corpus (6)
- Recovery Guarantees for One-hidden-layer Neural Networks
- Learning One-hidden-layer Neural Networks with Landscape Design
- On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
- The Landscape of Empirical Risk for Non-convex Losses
- Scaling Limit: Exact and Tractable Analysis of Online Learning Algorithms with Applications to Regularized Regression and PCA
- Provable Methods for Training Neural Networks with Sparse Connectivity
Cited by in corpus (213)
- Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Gradient Descent Finds Global Minima of Deep Neural Networks
- Analysis of the Generalization Error: Empirical Risk Minimization over Deep Artificial Neural Networks Overcomes the Curse of Dimensionality in the Numerical Approximation of Black-Scholes Partial Differential Equations
- Mean Field Limit for Coulomb-Type Flows
- Neural Network Approximation: Three Hidden Layers Are Enough
- Rademacher Complexity for Adversarially Robust Generalization
- Modelling the influence of data structure on learning in neural networks: the hidden manifold model
- A Priori Estimates of the Population Risk for Two-layer Neural Networks
- On the Convergence Rate of Training Recurrent Neural Networks
- How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
- A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics
- Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
- Exploring Deep Neural Networks via Layer-Peeled Model: Minority Collapse in Imbalanced Training
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- The Pitfalls of Simplicity Bias in Neural Networks
- Supervised learning from noisy observations: Combining machine-learning techniques with data assimilation
- When Do Neural Networks Outperform Kernel Methods?
- Theory of the Frequency Principle for General Deep Neural Networks
- Propagation of chaos: a review of models, methods and applications. II. Applications
- Hierarchies, entropy, and quantitative propagation of chaos for mean field diffusions
- The large learning rate phase of deep learning: the catapult mechanism
- Analyzing Upper Bounds on Mean Absolute Errors for Deep Neural Network Based Vector-to-Vector Regression
- On the Global Convergence of Particle Swarm Optimization Methods
- Kernel and Rich Regimes in Overparametrized Models
- Trainability and Accuracy of Neural Networks: An Interacting Particle System Approach
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Good Subnetworks Provably Exist: Pruning via Greedy Forward Selection
- Two-Layer Neural Networks for Partial Differential Equations: Optimization and Generalization Theory
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- Machine Learning from a Continuous Viewpoint
- The Gaussian equivalence of generative models for learning with shallow neural networks
- Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound
- The effective noise of Stochastic Gradient Descent
- A Selective Overview of Deep Learning
- A mean-field limit for certain deep neural networks
- Learning nonequilibrium control forces to characterize dynamical phase transitions
- Gradient Dynamics of Shallow Univariate ReLU Networks
- Mean-field inference methods for neural networks
- On the Expressive Power of Deep Polynomial Neural Networks
- Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
- Uncertainty Quantification and Deep Ensembles
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Mean-Field Langevin Dynamics and Energy Landscape of Neural Networks
- Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed
- Consensus-Based Optimization Methods Converge Globally
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
- Robustness of Bayesian Neural Networks to Gradient-Based Attacks
- Mathematical Models of Overparameterized Neural Networks
- Representation formulas and pointwise properties for Barron functions
- Geometric compression of invariant manifolds in neural nets
- Linear Frequency Principle Model to Understand the Absence of Overfitting in Neural Networks
- A Priori Generalization Analysis of the Deep Ritz Method for Solving High Dimensional Elliptic Equations
- Gradient Descent can Learn Less Over-parameterized Two-layer Neural Networks on Classification Problems
- Feature Learning in Infinite-Width Neural Networks
- Splitting Steepest Descent for Growing Neural Architectures
- Network size and weights size for memorization with two-layers neural networks
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- Beyond the storage capacity: data driven satisfiability transition
- PSO-Convolutional Neural Networks with Heterogeneous Learning Rate
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
- Phase diagram for two-layer ReLU neural networks at infinite-width limit
- Mean-Field Neural ODEs via Relaxed Optimal Control
- Sinkformers: Transformers with Doubly Stochastic Attention
- On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
- On the infinite width limit of neural networks with a standard parameterization
- Online Stochastic Gradient Descent with Arbitrary Initialization Solves Non-smooth, Non-convex Phase Retrieval
- Taylorized Training: Towards Better Approximation of Neural Network Training at Finite Width
- Quantitative Propagation of Chaos for SGD in Wide Neural Networks
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks
- EPR-Net: Constructing non-equilibrium potential landscape via a variational force projection formulation
- Reliable Off-policy Evaluation for Reinforcement Learning
- Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- Uniform-in-time propagation of chaos for kinetic mean field Langevin dynamics
- Mean-field Langevin System, Optimal Control and Deep Neural Networks
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- Deep Equals Shallow for ReLU Networks in Kernel Regimes
- Large-time asymptotics in deep learning
- The Local Elasticity of Neural Networks
- Coding schemes in neural networks learning classification tasks
- Classification Logit Two-sample Testing by Neural Networks
- Generalization and Memorization: The Bias Potential Model
- A mean-field analysis of two-player zero-sum games
- Overfitting Can Be Harmless for Basis Pursuit, But Only to a Degree
- Stationary Points of Shallow Neural Networks with Quadratic Activation Function
- Generalisation dynamics of online learning in over-parameterised neural networks
- A Local Convergence Theory for Mildly Over-Parameterized Two-Layer Neural Network
- On Sparsity in Overparametrised Shallow ReLU Networks
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification
- Mehler's Formula, Branching Process, and Compositional Kernels of Deep Neural Networks
- Propagation of chaos: a review of models, methods and applications. I. Models and methods
- Solving PDEs on Unknown Manifolds with Machine Learning
- Analysis of feature learning in weight-tied autoencoders via the mean field lens
- Adaptive and Implicit Regularization for Matrix Completion
- Spatially heterogeneous learning by a deep student machine
- AIR-Net: Adaptive and Implicit Regularization Neural Network for Matrix Completion
- Context Aware Machine Learning
- The asymptotic spectrum of the Hessian of DNN throughout training
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- Global Convergence of Second-order Dynamics in Two-layer Neural Networks
- The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models
- Polyak-Łojasiewicz inequality on the space of measures and convergence of mean-field birth-death processes
- Benign Overfitting and Noisy Features
- Unbiased deep solvers for linear parametric PDEs
- Maximum likelihood estimation of potential energy in interacting particle systems from single-trajectory data
- Theory III: Dynamics and Generalization in Deep Networks
- Non-asymptotic approximations of neural networks by Gaussian processes
- How Implicit Regularization of ReLU Neural Networks Characterizes the Learned Function -- Part I: the 1-D Case of Two Layers with Random First Layer
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural Networks
- A Spectral Analysis of Dot-product Kernels
- Learning Deep ReLU Networks Is Fixed-Parameter Tractable
- On Energy-Based Models with Overparametrized Shallow Neural Networks
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective
- How neural networks find generalizable solutions: Self-tuned annealing in deep learning
- McKean-Vlasov equations involving hitting times: blow-ups and global solvability
- Learning time-scales in two-layers neural networks
- Towards an Understanding of Residual Networks Using Neural Tangent Hierarchy (NTH)
- A Deep Conditioning Treatment of Neural Networks
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- A unified Fourier slice method to derive ridgelet transform for a variety of depth-2 neural networks
- A Measure Theoretical Approach to the Mean-field Maximum Principle for Training NeurODEs
- Global Convergence of Gradient Descent for Deep Linear Residual Networks
- Inference in Multi-Layer Networks with Matrix-Valued Unknowns
- Universal scaling laws in the gradient descent training of neural networks
- Soft Mode in the Dynamics of Over-realizable On-line Learning for Soft Committee Machines
- The Discovery of Dynamics via Linear Multistep Methods and Deep Learning: Error Estimation
- A Recipe for Global Convergence Guarantee in Deep Neural Networks
- Learning with Gradient Descent and Weakly Convex Losses
- Large deviations of one-hidden-layer neural networks
- A Diffusion Approximation Theory of Momentum SGD in Nonconvex Optimization
- SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Plateau Phenomenon in Gradient Descent Training of ReLU networks: Explanation, Quantification and Avoidance
- Prediction intervals for Deep Neural Networks
- KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support
- Well-posedness and approximation of reflected McKean-Vlasov SDEs with applications
- Sub-Optimal Local Minima Exist for Neural Networks with Almost All Non-Linear Activations
- Law of large numbers and central limit theorem for wide two-layer neural networks: the mini-batch and noisy case
- Uniform-in-time propagation of chaos for mean field Langevin dynamics
- On the Global Convergence of Gradient Descent for multi-layer ResNets in the mean-field regime
- Gradient flows on graphons: existence, convergence, continuity equations
- Why Lottery Ticket Wins? A Theoretical Perspective of Sample Complexity on Pruned Neural Networks
- Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
- The loss landscape of deep linear neural networks: a second-order analysis
- The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
- Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient Descent
- Scaling Neural Tangent Kernels via Sketching and Random Features
- Why Shallow Networks Struggle to Approximate and Learn High Frequencies
- Embedding Principle of Loss Landscape of Deep Neural Networks
- Particle Dual Averaging: Optimization of Mean Field Neural Networks with Global Convergence Rate Analysis
- Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
- Ergodicity of the underdamped mean-field Langevin dynamics
- Ridge Regression with Over-Parametrized Two-Layer Networks Converge to Ridgelet Spectrum
- An analytic theory of shallow networks dynamics for hinge loss classification
- Subaging in underparametrized Deep Neural Networks
- Dynamically Stable Infinite-Width Limits of Neural Classifiers
- SGD Distributional Dynamics of Three Layer Neural Networks
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- On functions computed on trees
- Global Convergence of SGD On Two Layer Neural Nets
- Phase Diagram of Initial Condensation for Two-layer Neural Networks
- Fast Approximation and Estimation Bounds of Kernel Quadrature for Infinitely Wide Models
- Implicit Bias of Linear RNNs
- Deep limits and cut-off phenomena for neural networks
- Neural Spectral Marked Point Processes
- Understanding Deflation Process in Over-parametrized Tensor Decomposition
- A Note on the Global Convergence of Multilayer Neural Networks in the Mean Field Regime
- Beyond Lazy Training for Over-parameterized Tensor Decomposition
- Do Input Gradients Highlight Discriminative Features?
- Experiments with Rich Regime Training for Deep Learning
- Towards a General Theory of Infinite-Width Limits of Neural Classifiers
- Dynamics of Meta-learning Representation in the Teacher-student Scenario
- Predicting the outputs of finite deep neural networks trained with noisy gradients
- Self-interacting approximation to McKean-Vlasov long-time limit: a Markov chain Monte Carlo method
- Mean-field limit of particle systems with absorption
- Learning One-hidden-layer neural networks via Provable Gradient Descent with Random Initialization
- Decomposed resolution of finite-state aggregative optimal control problems
- On the convergence of gradient descent for two layer neural networks
- Random Features for the Neural Tangent Kernel
- Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference
- Mirror Descent Algorithms for Minimizing Interacting Free Energy
- WGAN with an Infinitely Wide Generator Has No Spurious Stationary Points
- Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks
- Tighter Sparse Approximation Bounds for ReLU Neural Networks
- Unintended Effects on Adaptive Learning Rate for Training Neural Network with Output Scale Change
- A Mean-Field Theory for Learning the Schönberg Measure of Radial Basis Functions
- Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
- One-pass Stochastic Gradient Descent in Overparametrized Two-layer Neural Networks
- Ghosts in Neural Networks: Existence, Structure and Role of Infinite-Dimensional Null Space
- Dual Training of Energy-Based Models with Overparametrized Shallow Neural Networks
- Transfer entropy and O-information to detect grokking in tensor network multi-class classification problems
- On the exponential ergodicity of the McKean-Vlasov SDE depending on a polynomial interaction
- Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
- A trajectorial approach to relative entropy dissipation of McKeanVlasov diffusions: gradient flows and HWBI inequalities
- Mean-field Analysis of Piecewise Linear Solutions for Wide ReLU Networks
- A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks
- Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks
- Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation
- Statistical mechanics of extensive-width Bayesian neural networks near interpolation
- Probability distribution in the Toda system: The singular route to a steady state
- DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths
- Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Stationary Density Estimation of Itô Diffusions Using Deep Learning
- Deterministic and Stochastic Frank-Wolfe Recursion on Probability Spaces
- Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime
- A note on regularised NTK dynamics with an application to PAC-Bayesian training
- Towards Understanding Learning in Neural Networks with Linear Teachers