7 papers
Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables
Masaaki Imaizumi, Masanori Koyama, Noboru Isobe +1
We use a mean-field-based transformer model to theoretically investigate how auxiliary variables, such as positional encoding, prevent mode collapse of self-attention mechanisms. T…
Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers
Noboru Isobe, Daisuke Inoue, Masaaki Imaizumi
Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically, and recent mean-field theor…
Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers
Mana Sakai, Masaaki Imaizumi
Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled w…
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
Tomomasa Hara, Hiroto Kurita, Masaaki Imaizumi +2
For constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach. This paper examines whether mean pooling actually works well in real mode…
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
Sota Nishiyama, Masaaki Imaizumi
Modern machine learning models are typically trained via multi-pass stochastic gradient descent (SGD) with small batch sizes, and understanding their dynamics in high dimensions is…
Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent
Shota Imai, Sota Nishiyama, Masaaki Imaizumi
The dynamics of gradient-based training in neural networks often exhibit nontrivial structures; hence, understanding them remains a central challenge in theoretical machine learnin…