Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
arXiv:2410.18858 · doi:10.1103/l4p2-vrxt
Abstract
Current progress in artificial intelligence is centered around so-called large language models that consist of neural networks processing long sequences of high-dimensional vectors called tokens. Statistical physics provides powerful tools to study the functioning of learning with neural networks and has played a recognized role in the development of modern machine learning. The statistical physics approach relies on simplified and analytically tractable models of data. However, simple tractable models for long sequences of high-dimensional tokens are largely underexplored. Inspired by the crucial role models such as the single-layer teacher-student perceptron (aka generalized linear regression) played in the theory of fully connected neural networks, in this paper, we introduce and study the bilinear sequence regression (BSR) as one of the most basic models for sequences of tokens. We note that modern architectures naturally subsume the BSR model due to the skip connections. Building on recent methodological progress, we compute the Bayes-optimal generalization error for the model in the limit of long sequences of high-dimensional tokens, and provide a message-passing algorithm that matches this performance. We quantify the improvement that optimal learning brings with respect to vectorizing the sequence of tokens and learning via simple linear regression. We also unveil surprising properties of the gradient descent algorithms in the BSR model.
References in corpus (21)
- Guaranteed Minimum-Rank Solutions of Linear Matrix Equations via Nuclear Norm Minimization
- Message Passing Algorithms for Compressed Sensing
- Reconciling modern machine learning practice and the bias-variance trade-off
- Statistical physics of inference: Thresholds and algorithms
- Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
- Statistical physics-based reconstruction in compressed sensing
- Multilinear tensor regression for longitudinal relational data
- Minimax risk of matrix denoising by singular value thresholding
- Subdominant Dense Clusters Allow for Simple Learning and High Computational Performance in Neural Networks with Discrete Synapses
- Modelling the influence of data structure on learning in neural networks: the hidden manifold model
- Statistical Inference for High-Dimensional Matrix-Variate Factor Model
- The Phase Transition of Matrix Recovery from Gaussian Measurements Matches the Minimax MSE of Matrix Denoising
- Mapping of attention mechanisms to a generalized Potts model
- Statistical limits of dictionary learning: random matrix theory and the spectral replica method
- Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising
- Matrix factorization with neural networks
- A spin glass model for reconstructing nonlinearly encrypted signals corrupted by noise
- Near-optimal matrix recovery from random linear measurements
- Matrix denoising: Bayes-optimal estimators via low-degree polynomials
- Phase diagram of matrix compressed sensing
- Dynamical mean field theory for models of confluent tissues and beyond