Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
arXiv:2402.02362 · doi:10.1088/2632-2153/ad5927
Abstract
Understanding the inner workings of neural networks, including transformers, remains one of the most challenging puzzles in machine learning. This study introduces a novel approach by applying the principles of gauge symmetries, a key concept in physics, to neural network architectures. By regarding model functions as physical observables, we find that parametric redundancies of various machine learning models can be interpreted as gauge symmetries. We mathematically formulate the parametric redundancies in neural ODEs, and find that their gauge symmetries are given by spacetime diffeomorphisms, which play a fundamental role in Einstein's theory of gravity. Viewing neural ODEs as a continuum version of feedforward neural networks, we show that the parametric redundancies in feedforward neural networks are indeed lifted to diffeomorphisms in neural ODEs. We further extend our analysis to transformer models, finding natural correspondences with neural ODEs and their gauge symmetries. The concept of gauge symmetries sheds light on the complex behavior of deep learning models through physics and provides us with a unifying perspective for analyzing various machine learning architectures.
11 pages, 3 figures
References in corpus (16)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- cosFormer: Rethinking Softmax in Attention
- AdS/CFT as a deep Boltzmann machine
- Deep Learning and AdS/QCD
- A Study on ReLU and Softmax in Transformer
- Learning the black hole metric from holographic conductivity
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Machine Learning Statistical Gravity from Multi-Region Entanglement Entropy
- Inferring effective couplings with Restricted Boltzmann Machines
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances
- Visualizing high-dimensional loss landscapes with Hessian directions
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
- On the Symmetries of Deep Learning Models and their Internal Representations
- A Neural ODE Interpretation of Transformer Layers