Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
arXiv:1312.6120
Abstract
Despite the widespread practical success of deep learning methods, our theoretical understanding of the dynamics of learning in deep neural networks remains quite sparse. We attempt to bridge the gap between the theory and practice of deep learning by systematically analyzing learning dynamics for the restricted case of deep linear neural networks. Despite the linearity of their input-output map, such networks have nonlinear gradient descent dynamics on weights that change with the addition of each new hidden layer. We show that deep linear networks exhibit nonlinear learning phenomena similar to those seen in simulations of nonlinear networks, including long plateaus followed by rapid transitions to lower error solutions, and faster convergence from greedy unsupervised pretraining initial conditions than from random initial conditions. We provide an analytical description of these phenomena by finding new exact solutions to the nonlinear dynamics of deep learning. Our theoretical analysis also reveals the surprising finding that as the depth of a network approaches infinity, learning speed can nevertheless remain finite: for a special class of initial conditions on the weights, very deep networks incur only a finite, depth independent, delay in learning speed relative to shallow networks. We show that, under certain conditions on the training data, unsupervised pretraining can find this special class of initial conditions, while scaled random Gaussian initializations cannot. We further exhibit a new class of random orthogonal initial conditions on weights that, like unsupervised pre-training, enjoys depth independent learning times. We further show that these initial conditions also lead to faithful propagation of gradients even in deep nonlinear networks, as long as they operate in a special regime known as the edge of chaos.
Submission to ICLR2014. Revised based on reviewer feedback
References in corpus (1)
Cited by in corpus (322)
- Deep Residual Learning for Image Recognition
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Training Very Deep Networks
- Activation Functions: Comparison of trends in Practice and Research for Deep Learning
- Circuit-centric quantum classifiers
- Applications of Deep Learning and Reinforcement Learning to Biological Data
- The Loss Surfaces of Multilayer Networks
- Resnet in Resnet: Generalizing Residual Architectures
- Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Artificial neural networks for neuroscientists: A primer
- A Closer Look at Memorization in Deep Networks
- Single-source Domain Expansion Network for Cross-Scene Hyperspectral Image Classification
- Working hard to know your neighbor's margins: Local descriptor learning loss
- Convergence Analysis of Two-layer Neural Networks with ReLU Activation
- Robust Large Margin Deep Neural Networks
- Doctor AI: Predicting Clinical Events via Recurrent Neural Networks
- The Principles of Deep Learning Theory
- Barren Plateaus in Variational Quantum Computing
- Star-galaxy Classification Using Deep Convolutional Neural Networks
- Qualitatively characterizing neural network optimization problems
- Data-dependent Initializations of Convolutional Neural Networks
- Regularization for Deep Learning: A Taxonomy
- Escaping From Saddle Points --- Online Stochastic Gradient for Tensor Decomposition
- Direct Feedback Alignment Provides Learning in Deep Neural Networks
- Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?
- Scaling description of generalization with number of parameters in deep learning
- On orthogonality and learning recurrent networks with long term dependencies
- ReZero is All You Need: Fast Convergence at Large Depth
- Adversarial Video Generation on Complex Datasets
- Checkerboard artifact free sub-pixel convolution: A note on sub-pixel convolution, resize convolution and convolution resize
- Entropy and mutual information in models of deep neural networks
- Sharp Minima Can Generalize For Deep Nets
- Optimization for deep learning: theory and algorithms
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
- Self-Adaptive Physics-Informed Neural Networks using a Soft Attention Mechanism
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Fixup Initialization: Residual Learning Without Normalization
- Gradient Descent Happens in a Tiny Subspace
- Overfitting Mechanism and Avoidance in Deep Neural Networks
- Review: Deep Learning in Electron Microscopy
- InSituNet: Deep Image Synthesis for Parameter Space Exploration of Ensemble Simulations
- Understanding the Role of Training Regimes in Continual Learning
- Tunable Efficient Unitary Neural Networks (EUNN) and their application to RNNs
- Universal discriminative quantum neural networks
- Multiplicative LSTM for sequence modelling
- Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- Regularizing CNNs with Locally Constrained Decorrelations
- Deep Polynomial Neural Networks
- Avoiding pathologies in very deep networks
- BBN: Bilateral-Branch Network with Cumulative Learning for Long-Tailed Visual Recognition
- NewsQA: A Machine Comprehension Dataset
- Towards an integration of deep learning and neuroscience
- On the saddle point problem for non-convex optimization
- On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping
- Random Walk Initialization for Training Very Deep Feedforward Networks
- An Analytical Formula of Population Gradient for two-layered ReLU network and its Applications in Convergence and Critical Point Analysis
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- Deep learning as a tool for neural data analysis: speech classification and cross-frequency coupling in human sensorimotor cortex
- Modeling Others using Oneself in Multi-Agent Reinforcement Learning
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- Three-dimensional microstructure generation using generative adversarial neural networks in the context of continuum micromechanics
- Analysis and Optimization of Convolutional Neural Network Architectures
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- Thermodynamics of Restricted Boltzmann Machines and related learning dynamics
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Newton-Type Methods for Non-Convex Optimization Under Inexact Hessian Information
- An analytic theory of generalization dynamics and transfer learning in deep linear networks
- What shapes feature representations? Exploring datasets, architectures, and training
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
- How to Start Training: The Effect of Initialization and Architecture
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- FedBABU: Towards Enhanced Representation for Federated Image Classification
- Three Mechanisms of Weight Decay Regularization
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- Spectral Dynamics of Learning Restricted Boltzmann Machines
- Effect of Depth and Width on Local Minima in Deep Learning
- Gradient Starvation: A Learning Proclivity in Neural Networks
- Improved training of binary networks for human pose estimation and image recognition
- Local minima in training of neural networks
- Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000-Layer Vanilla Convolutional Neural Networks
- Portraying Double Higgs at the Large Hadron Collider
- The worst of both worlds: A comparative analysis of errors in learning from data in psychology and machine learning
- On2Vec: Embedding-based Relation Prediction for Ontology Population
- Convolutional Neural Networks Analyzed via Convolutional Sparse Coding
- Learning to Discover Efficient Mathematical Identities
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- Principled Weight Initialization for Hypernetworks
- Consensus Attention-based Neural Networks for Chinese Reading Comprehension
- Second-Order Optimization for Non-Convex Machine Learning: An Empirical Study
- A Comprehensive and Modularized Statistical Framework for Gradient Norm Equality in Deep Neural Networks
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- The Perfect Match: 3D Point Cloud Matching with Smoothed Densities
- Fundamental bounds on learning performance in neural circuits
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Gated Recurrent Neural Tensor Network
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- Data-driven emergence of convolutional structure in neural networks
- How far can we go without convolution: Improving fully-connected networks
- Inexact Non-Convex Newton-Type Methods
- Why Spectral Normalization Stabilizes GANs: Analysis and Improvements
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Learning through atypical "phase transitions" in overparameterized neural networks
- Understanding the Difficulty of Training Transformers
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Toward Continual Learning for Conversational Agents
- Joint Sequence Learning and Cross-Modality Convolution for 3D Biomedical Segmentation
- Two Routes to Scalable Credit Assignment without Weight Symmetry
- Learning Implicitly Recurrent CNNs Through Parameter Sharing
- Dynamical Isometry and a Mean Field Theory of LSTMs and GRUs
- Ultra-High-Resolution Detector Simulation with Intra-Event Aware GAN and Self-Supervised Relational Reasoning
- Regularization Methods for Generative Adversarial Networks: An Overview of Recent Studies
- Exponential Convergence Time of Gradient Descent for One-Dimensional Deep Linear Neural Networks
- Data-Efficient Learning of Feedback Policies from Image Pixels using Deep Dynamical Models
- Active Long Term Memory Networks
- Geodesics of learned representations
- DizzyRNN: Reparameterizing Recurrent Neural Networks for Norm-Preserving Backpropagation
- A Spectral Energy Distance for Parallel Speech Synthesis
- Beyond Shared Hierarchies: Deep Multitask Learning through Soft Layer Ordering
- On Generalization Bounds of a Family of Recurrent Neural Networks
- Training Agents using Upside-Down Reinforcement Learning
- Online Learning with Gated Linear Networks
- MinimalRNN: Toward More Interpretable and Trainable Recurrent Neural Networks
- Finite size corrections for neural network Gaussian processes
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Utilizing the Instability in Weakly Supervised Object Detection
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- Can recurrent neural networks warp time?
- On Data-Augmentation and Consistency-Based Semi-Supervised Learning
- Depth induces scale-averaging in overparameterized linear Bayesian neural networks
- GAMI-Net: An Explainable Neural Network based on Generalized Additive Models with Structured Interactions
- Learning to Represent Words in Context with Multilingual Supervision
- Sparse Deep Neural Network Exact Solutions
- Degeneration in VAE: in the Light of Fisher Information Loss
- Multi-center validation study of automated classification of pathological slowing in adult scalp electroencephalograms via frequency features
- Jointly Predicting Predicates and Arguments in Neural Semantic Role Labeling
- Singular Value Decomposition and Neural Networks
- Transforming task representations to perform novel tasks
- Bio-JOIE: Joint Representation Learning of Biological Knowledge Bases
- Pre-Trained Models: Past, Present and Future
- Attributed Sequence Embedding
- Optimization and Generalization of Regularization-Based Continual Learning: a Loss Approximation Viewpoint
- Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?
- The Benefits of Over-parameterization at Initialization in Deep ReLU Networks
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- Revealing the Structure of Deep Neural Networks via Convex Duality
- Deep Residual Networks and Weight Initialization
- On the geometry of generalization and memorization in deep neural networks
- ReFine: Re-randomization before Fine-tuning for Cross-domain Few-shot Learning
- Denoising without access to clean data using a partitioned autoencoder
- On the energy landscape of deep networks
- Residual Tensor Train: A Quantum-inspired Approach for Learning Multiple Multilinear Correlations
- Two-phase flow regime prediction using LSTM based deep recurrent neural network
- On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization
- CCLF: A Contrastive-Curiosity-Driven Learning Framework for Sample-Efficient Reinforcement Learning
- The Implicit Bias of Depth: How Incremental Learning Drives Generalization
- Non-Gaussian processes and neural networks at finite widths
- Orthogonalizing Convolutional Layers with the Cayley Transform
- When Does Preconditioning Help or Hurt Generalization?
- The Low-Rank Simplicity Bias in Deep Networks
- Inference in Deep Networks in High Dimensions
- Emergence of Network Motifs in Deep Neural Networks
- A Weight Initialization Based on the Linear Product Structure for Neural Networks
- Convolutional Residual Memory Networks
- Norm-preserving Orthogonal Permutation Linear Unit Activation Functions (OPLU)
- XIRL: Cross-embodiment Inverse Reinforcement Learning
- On the Stability of Deep Networks
- High Dimensional Classification via Regularized and Unregularized Empirical Risk Minimization: Precise Error and Optimal Loss
- Towards Robust Deep Neural Networks
- Cascade EF-GAN: Progressive Facial Expression Editing with Local Focuses
- Estimating entropy production in a stochastic system with odd-parity variables
- Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
- Feature Flow Regularization: Improving Structured Sparsity in Deep Neural Networks
- SAT-NGP : Unleashing Neural Graphics Primitives for Fast Relightable Transient-Free 3D reconstruction from Satellite Imagery
- An Empirical Exploration of Skip Connections for Sequential Tagging
- On Dropout and Nuclear Norm Regularization
- Mutual Information Scaling and Expressive Power of Sequence Models
- Batch Normalization Provably Avoids Rank Collapse for Randomly Initialised Deep Networks
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- A Unified Paths Perspective for Pruning at Initialization
- Breaking the Activation Function Bottleneck through Adaptive Parameterization
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Regularizing activations in neural networks via distribution matching with the Wasserstein metric
- Biological credit assignment through dynamic inversion of feedforward networks
- Deep linear neural networks with arbitrary loss: All local minima are global
- Implicit Regularization of Stochastic Gradient Descent in Natural Language Processing: Observations and Implications
- TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations
- A deep learning theory for neural networks grounded in physics
- Fluctuation-dissipation Type Theorem in Stochastic Linear Learning
- Fast Face-swap Using Convolutional Neural Networks
- Information Geometry of Orthogonal Initializations and Training
- How to Make Deep RL Work in Practice
- Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
- Convolutional Normalization: Improving Deep Convolutional Network Robustness and Training
- Neuron Campaign for Initialization Guided by Information Bottleneck Theory
- Analysis of feature learning in weight-tied autoencoders via the mean field lens
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification
- Neural networks for semantic segmentation of historical city maps: Cross-cultural performance and the impact of figurative diversity
- AMEIR: Automatic Behavior Modeling, Interaction Exploration and MLP Investigation in the Recommender System
- Predicting Training Time Without Training
- Fusion Graph Convolutional Networks
- Implicit Regularization via Neural Feature Alignment
- Fast Certified Robust Training with Short Warmup
- Inversion dynamics of class manifolds in deep learning reveals tradeoffs underlying generalisation
- Information Bottleneck Theory on Convolutional Neural Networks
- Multi-Level Recurrent Residual Networks for Action Recognition
- Constraint-Based Regularization of Neural Networks
- Convolution Aware Initialization
- A Sample Complexity Separation between Non-Convex and Convex Meta-Learning
- Numerically Recovering the Critical Points of a Deep Linear Autoencoder
- A global convergence theory for deep ReLU implicit networks via over-parameterization
- Towards Understanding the Generalization Bias of Two Layer Convolutional Linear Classifiers with Gradient Descent
- Deep Learning for Inverse Problems: Bounds and Regularizers
- Measuring and Understanding Sensory Representations within Deep Networks Using a Numerical Optimization Framework
- RHR-Net: A Residual Hourglass Recurrent Neural Network for Speech Enhancement
- Proof-of-Learning: Definitions and Practice
- Spectrum concentration in deep residual learning: a free probability approach
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- Analysis on Gradient Propagation in Batch Normalized Residual Networks
- Why bigger is not always better: on finite and infinite neural networks
- Deep frequency principle towards understanding why deeper learning is faster
- Multi-Agent Deep Reinforcement Learning for Large-scale Traffic Signal Control
- Learning Dynamics of Linear Denoising Autoencoders
- An Adversarial Transfer Network for Knowledge Representation Learning
- A learning gap between neuroscience and reinforcement learning
- Instance-Based Learning of Span Representations: A Case Study through Named Entity Recognition
- Deep Probabilistic Time Series Forecasting using Augmented Recurrent Input for Dynamic Systems
- A Span Selection Model for Semantic Role Labeling
- Global Convergence of Gradient Descent for Deep Linear Residual Networks
- Auxiliary Learning by Implicit Differentiation
- DeepFolio: Convolutional Neural Networks for Portfolios with Limit Order Book Data
- A Geometric Approach of Gradient Descent Algorithms in Linear Neural Networks
- SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
- Audio-Conditioned U-Net for Position Estimation in Full Sheet Images
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Understanding Dynamics of Nonlinear Representation Learning and Its Application
- DCT-Conv: Coding filters in convolutional networks with Discrete Cosine Transform
- Escaping spurious local minimum trajectories in online time-varying nonconvex optimization
- Regularizing Neural Networks via Minimizing Hyperspherical Energy
- Collective evolution of weights in wide neural networks
- Learning to Read and Follow Music in Complete Score Sheet Images
- Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNs
- Discovery of Latent Factors in High-dimensional Data Using Tensor Methods
- GradNets: Dynamic Interpolation Between Neural Architectures
- Syntax-based Attention Model for Natural Language Inference
- Learn Faster and Forget Slower via Fast and Stable Task Adaptation
- Shortcut Sequence Tagging
- Implicit Sparse Regularization: The Impact of Depth and Early Stopping
- Fractional moment-preserving initialization schemes for training deep neural networks
- Reborn Mechanism: Rethinking the Negative Phase Information Flow in Convolutional Neural Network
- Trap of Feature Diversity in the Learning of MLPs
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- The loss landscape of deep linear neural networks: a second-order analysis
- Orthogonal Wasserstein GANs
- The Three Stages of Learning Dynamics in High-Dimensional Kernel Methods
- Grounded and Controllable Image Completion by Incorporating Lexical Semantics
- Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
- Residual Networks as Nonlinear Systems: Stability Analysis using Linearization
- Statistical Parametric Speech Synthesis Using Bottleneck Representation From Sequence Auto-encoder
- Connecting Graph Convolutional Networks and Graph-Regularized PCA
- Adjoined Networks: A Training Paradigm with Applications to Network Compression
- Are Saddles Good Enough for Deep Learning?
- Is BERT a Cross-Disciplinary Knowledge Learner? A Surprising Finding of Pre-trained Models' Transferability
- On Optimality Conditions for Auto-Encoder Signal Recovery
- Data-driven Weight Initialization with Sylvester Solvers
- Beyond Folklore: A Scaling Calculus for the Design and Initialization of ReLU Networks
- Neural Bayes: A Generic Parameterization Method for Unsupervised Representation Learning
- Improving Label Quality by Jointly Modeling Items and Annotators
- Summary statistics of learning link changing neural representations to behavior
- On the validity of kernel approximations for orthogonally-initialized neural networks
- A Newton-Based Method for Nonconvex Optimization with Fast Evasion of Saddle Points
- Neural Networks as Kernel Learners: The Silent Alignment Effect
- Residual Networks: Lyapunov Stability and Convex Decomposition
- Propagate-Selector: Detecting Supporting Sentences for Question Answering via Graph Neural Networks
- On regularization of gradient descent, layer imbalance and flat minima
- Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
- NeuroFabric: Identifying Ideal Topologies for Training A Priori Sparse Networks
- Universal Representation Learning of Knowledge Bases by Jointly Embedding Instances and Ontological Concepts
- Understanding Deflation Process in Over-parametrized Tensor Decomposition
- Mixed Moments for the Product of Ginibre Matrices
- MLAS: Metric Learning on Attributed Sequences
- Mack-Net model: Blending Mack's model with Recurrent Neural Networks
- Layer-Wise Interpretation of Deep Neural Networks Using Identity Initialization
- Exploiting Invertible Decoders for Unsupervised Sentence Representation Learning
- Notes on Deep Learning Theory
- Single Model Ensemble using Pseudo-Tags and Distinct Vectors
- Geometry Perspective Of Estimating Learning Capability Of Neural Networks
- Deep orthogonal linear networks are shallow
- Training Linear Neural Networks: Non-Local Convergence and Complexity Results
- NeurIPS 2019 Disentanglement Challenge: Improved Disentanglement through Learned Aggregation of Convolutional Feature Maps
- Being curious about the answers to questions: novelty search with learned attention
- Step Size Matters in Deep Learning
- A Study of Policy Gradient on a Class of Exactly Solvable Models
- Span-based discontinuous constituency parsing: a family of exact chart-based algorithms with time complexities from O(n^6) down to O(n^3)
- Post-Workshop Report on Science meets Engineering in Deep Learning, NeurIPS 2019, Vancouver
- A priori guarantees of finite-time convergence for Deep Neural Networks
- Predicting the success of Gradient Descent for a particular Dataset-Architecture-Initialization (DAI)
- SHORING: Design Provable Conditional High-Order Interaction Network via Symbolic Testing
- Neural Path Features and Neural Path Kernel : Understanding the role of gates in deep learning
- Medi-Care AI: Predicting Medications From Billing Codes via Robust Recurrent Neural Networks
- Diversity Regularized Adversarial Learning
- Feature Tracking Cardiac Magnetic Resonance via Deep Learning and Spline Optimization
- Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks
- Generating the support with extreme value losses
- Understanding the wiring evolution in differentiable neural architecture search
- An Empirical Evaluation Study on the Training of SDC Features for Dense Pixel Matching
- Layer Dynamics of Linearised Neural Nets
- THG: Transformer with Hyperbolic Geometry
- Learning Longer-term Dependencies via Grouped Distributor Unit
- An Effective Training Method For Deep Convolutional Neural Network
- Training Efficiency and Robustness in Deep Learning
- How regularization affects the critical points in linear networks
- The staircase property: How hierarchical structure can guide deep learning
- A Johnson--Lindenstrauss Framework for Randomly Initialized CNNs
- Large-Dimensional Random Matrix Theory and Its Applications in Deep Learning and Wireless Communications
- Student-Teacher Learning from Clean Inputs to Noisy Inputs
- The Landscape of Multi-Layer Linear Neural Network From the Perspective of Algebraic Geometry