7 papers
Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models
Ken Takeda, Masafumi Oizumi, Ryo Karakida
Generative models, including diffusion models, are increasingly used as foundation models and adapted through sequential fine-tuning, making continual learning an essential problem…
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
Tomohiro Hayase, Ryo Karakida
Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the con…
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
Tomohiro Hayase, Benoît Collins, Ryo Karakida
Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective…
The Impact of Anisotropic Covariance Structure on the Training Dynamics and Generalization Error of Linear Networks
Taishi Watanabe, Ryo Karakida, Jun-nosuke Teramae
The success of deep neural networks largely depends on the statistical structure of the training data. While learning dynamics and generalization on isotropic data are well-establi…
Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
Mana Sakai, Ryo Karakida, Masaaki Imaizumi
In modern theoretical analyses of neural networks, the infinite-width limit is often invoked to justify Gaussian approximations of neuron preactivations (e.g., via neural network G…
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
Akiyoshi Tomihari, Ryo Karakida
The theoretical understanding of self-attention (SA) has been steadily progressing. A prominent line of work studies a class of SA layers that admit an energy function decreased by…