A mathematical theory of semantic development in deep neural networks
arXiv:1810.10531 · doi:10.1073/pnas.1820226116
Abstract
An extensive body of empirical research has revealed remarkable regularities in the acquisition, organization, deployment, and neural representation of human semantic knowledge, thereby raising a fundamental conceptual question: what are the theoretical principles governing the ability of neural networks to acquire, organize, and deploy abstract knowledge by integrating across many individual experiences? We address this question by mathematically analyzing the nonlinear dynamics of learning in deep linear networks. We find exact solutions to this learning dynamics that yield a conceptual explanation for the prevalence of many disparate phenomena in semantic cognition, including the hierarchical differentiation of concepts through rapid developmental transitions, the ubiquity of semantic illusions between such transitions, the emergence of item typicality and category coherence as factors controlling the speed of semantic processing, changing patterns of inductive projection over development, and the conservation of semantic similarity in neural representations across species. Thus, surprisingly, our simple neural model qualitatively recapitulates many diverse regularities underlying semantic development, while providing analytic insight into how the statistical structure of an environment can interact with nonlinear deep learning dynamics to give rise to these regularities.
Cited by in corpus (49)
- Artificial neural networks for neuroscientists: A primer
- Understanding Dimensional Collapse in Contrastive Self-supervised Learning
- Few-Shot Adaptation of Generative Adversarial Networks
- Statistical Mechanics of Deep Linear Neural Networks: The Back-Propagating Kernel Renormalization
- What shapes feature representations? Exploring datasets, architectures, and training
- A neural network walks into a lab: towards using deep nets as models for human behavior
- Gradient Starvation: A Learning Proclivity in Neural Networks
- Universality and individuality in neural dynamics across large populations of recurrent networks
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit
- Understanding Self-supervised Learning with Dual Deep Networks
- Data-driven emergence of convolutional structure in neural networks
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
- How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
- How Do Recommendation Models Amplify Popularity Bias? An Analysis from the Spectral Perspective
- Artificial selection of communities drives the emergence of structured interactions
- When MAML Can Adapt Fast and How to Assist When It Cannot
- Linear Classification of Neural Manifolds with Correlated Variability
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- Three Learning Stages and Accuracy-Efficiency Tradeoff of Restricted Boltzmann Machines
- Decomposing neural networks as mappings of correlation functions
- Implicit Rank-Minimizing Autoencoder
- The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
- Regularized linear autoencoders recover the principal components, eventually
- Emergence of Network Motifs in Deep Neural Networks
- Singular Vectors of Sums of Rectangular Random Matrices and Optimal Estimators of High-Rank Signals: The Extensive Spike Model
- Fluctuation-dissipation Type Theorem in Stochastic Linear Learning
- If deep learning is the answer, then what is the question?
- Topology, Vorticity and Limit Cycle in a Stabilized Kuramoto-Sivashinsky Equation
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- A Sample Complexity Separation between Non-Convex and Convex Meta-Learning
- Let's Agree to Agree: Neural Networks Share Classification Order on Real Datasets
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Abstraction Mechanisms Predict Generalization in Deep Neural Networks
- The loss landscape of deep linear neural networks: a second-order analysis
- Weight fluctuations in (deep) linear neural networks and a derivation of the inverse-variance flatness relation
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Meta-Learning Strategies through Value Maximization in Neural Networks
- Memory and attention in deep learning
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations
- Word Interdependence Exposes How LSTMs Compose Representations
- Task Guided Compositional Representation Learning for ZDA
- Ordinal Characterization of Similarity Judgments
- Robustness of the Random Language Model
- LSTMs Compose (and Learn) Bottom-Up
- An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
- A Mechanism for Producing Aligned Latent Spaces with Autoencoders
- Post-Workshop Report on Science meets Engineering in Deep Learning, NeurIPS 2019, Vancouver