Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
Tomohiro Hayase, Ryo Karakida
Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the con…
stat.ML2026
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
Tomohiro Hayase, Benoît Collins, Ryo Karakida
Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective…
stat.ML2026
The Impact of Anisotropic Covariance Structure on the Training Dynamics and Generalization Error of Linear Networks
Taishi Watanabe, Ryo Karakida, Jun-nosuke Teramae
The success of deep neural networks largely depends on the statistical structure of the training data. While learning dynamics and generalization on isotropic data are well-establi…