6 papers
Reachability and asymptotics of Gaussian Transformer dynamics
Albert Alcalde, Zhengping Ji, Enrique Zuazua
We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability m…
Exact Sequence Interpolation with Transformers
Albert Alcalde, Giovanni Fantuzzi, Enrique Zuazua
We prove that transformers can exactly interpolate datasets of finite input sequences in , , with corresponding output sequences of smaller or equal length.…
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
Albert Alcalde, Giovanni Fantuzzi, Enrique Zuazua
Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers…
Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
Albert Alcalde, Leon Bungert, Konstantin Riedl +1
Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the e…
Time-Delayed Transformers for Data-Driven Modeling of Low-Dimensional Dynamics
Albert Alcalde, Markus Widhalm, Emre Yılmaz
We propose the time-delayed transformer (TD-TF), a simplified transformer architecture for data-driven modeling of unsteady spatio-temporal dynamics. TD-TF bridges linear operator-…
Attention's forward pass and Frank-Wolfe
Albert Alcalde, Borjan Geshkovski, Domènec Ruiz-Balet
We study the hardmax limit of self-attention dynamics for token embeddings obtained in the zero-temperature () regime, and relate it to the finite- setting. In th…