Mapping of attention mechanisms to a generalized Potts model
arXiv:2304.07235 · doi:10.1103/PhysRevResearch.6.023057
Abstract
Transformers are neural networks that revolutionized natural language processing and machine learning. They process sequences of inputs, like words, using a mechanism called self-attention, which is trained via masked language modeling (MLM). In MLM, a word is randomly masked in an input sequence, and the network is trained to predict the missing word. Despite the practical success of transformers, it remains unclear what type of data distribution self-attention can learn efficiently. Here, we show analytically that if one decouples the treatment of word positions and embeddings, a single layer of self-attention learns the conditionals of a generalized Potts model with interactions between sites and Potts colors. Moreover, we show that training this neural network is exactly equivalent to solving the inverse Potts problem by the so-called pseudo-likelihood method, well known in statistical physics. Using this mapping, we compute the generalization error of self-attention in a model scenario analytically using the replica method.
5 pages, 3 figures
References in corpus (6)
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- High-dimensional Ising model selection using -regularized logistic regression
- Introduction to Random Matrices - Theory and Practice
- Transformer variational wave functions for frustrated quantum spin systems
- Data-driven emergence of convolutional structure in neural networks
- The impact of memory on learning sequence-to-sequence tasks
Cited by in corpus (14)
- A simple linear algebra identity to optimize Large-Scale Neural Network Quantum States
- Transformer Wave Function for two dimensional frustrated magnets: emergence of a Spin-Liquid Phase in the Shastry-Sutherland Model
- Transformer neural networks and quantum simulators: a hybrid approach for simulating strongly correlated systems
- Eight challenges in developing theory of intelligence
- Fine-tuning Neural Network Quantum States
- Are queries and keys always relevant? A case study on Transformer wave functions
- Time-dependent Neural Galerkin Method for Quantum Dynamics
- Design principles of deep translationally-symmetric neural quantum states for frustrated magnets
- Physics-informed Transformers for Electronic Quantum States
- Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
- A simplified Parisi Ansatz II: REM universality
- Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
- Universal scaling limits for spin networks via martingale methods
- Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings