2 papers
cs.LG2025
Multistability of Self-Attention Dynamics in Transformers
Claudio Altafini
In machine learning, a self-attention dynamics is a continuous-time multiagent-like model of the attention mechanisms of transformers. In this paper we show that such dynamics is r…
cs.LG2025
Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective
Joel Wendin, Claudio Altafini
The paper surveys recent progresses in understanding the dynamics and loss landscape of the gradient flow equations associated to deep linear neural networks, i.e., the gradient de…