Network Dynamics-Based Framework for Understanding Deep Neural Networks
arXiv:2501.02436 · doi:10.1007/s11433-025-2929-5
Abstract
Advancements in artificial intelligence call for a deeper understanding of the fundamental mechanisms underlying deep learning. In this work, we propose a theoretical framework to analyze learning dynamics through the lens of dynamical systems theory. We redefine the notions of linearity and nonlinearity in neural networks by introducing two fundamental transformation units at the neuron level: order-preserving transformations and non-order-preserving transformations. Different transformation modes lead to distinct collective behaviors in weight vector organization, different modes of information extraction, and the emergence of qualitatively different learning phases. Transitions between these phases may occur during training, accounting for key phenomena such as grokking. To further characterize generalization and structural stability, we introduce the concept of attraction basins in both sample and weight spaces. The distribution of neurons with different transformation modes across layers, along with the structural characteristics of the two types of attraction basins, forms a set of core metrics for analyzing the performance of learning models. Hyperparameters such as depth, width, learning rate, and batch size act as control variables for fine-tuning these metrics. Our framework not only sheds light on the intrinsic advantages of deep learning, but also provides a novel perspective for optimizing network architectures and training strategies.
13 pages, 7 figures
References in corpus (15)
- Reconciling modern machine learning practice and the bias-variance trade-off
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Unreasonable Effectiveness of Learning Neural Networks: From Accessible States and Robust Ensembles to Basic Algorithmic Schemes
- Explaining Neural Scaling Laws
- Supervised Learning with Projected Entangled Pair States
- Disentangling feature and lazy training in deep neural networks
- Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions
- Neural network interpretation using descrambler groups
- Inferring Global Dynamics of a Black-Box System Using Machine Learning
- Inferring Global Dynamics Using a Learning Machine
- Copy the dynamics using a learning machine
- Finite-time Lyapunov exponents of deep neural networks
- Cycle-tree guided attack of random K-core: Spin glass model and efficient message-passing algorithm
- Energy--Information Trade-off Induces Continuous and Discontinuous Phase Transitions in Lateral Predictive Coding
- Fermi-Bose Machine achieves both generalization and adversarial robustness