collaborators

6 papers

cs.LG2026

Reachability and asymptotics of Gaussian Transformer dynamics

Albert Alcalde, Zhengping Ji, Enrique Zuazua

We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability m…

cs.LG2026

Exact Sequence Interpolation with Transformers

Albert Alcalde, Giovanni Fantuzzi, Enrique Zuazua

We prove that transformers can exactly interpolate datasets of finite input sequences in , , with corresponding output sequences of smaller or equal length.…

cs.CL2026

Clustering in pure-attention hardmax transformers and its role in sentiment analysis

Albert Alcalde, Giovanni Fantuzzi, Enrique Zuazua

Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers…

math.AP2026

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

Albert Alcalde, Leon Bungert, Konstantin Riedl +1

Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the e…

cs.LG2026

Time-Delayed Transformers for Data-Driven Modeling of Low-Dimensional Dynamics

Albert Alcalde, Markus Widhalm, Emre Yılmaz

We propose the time-delayed transformer (TD-TF), a simplified transformer architecture for data-driven modeling of unsteady spatio-temporal dynamics. TD-TF bridges linear operator-…

math.OC2025

Attention's forward pass and Frank-Wolfe

Albert Alcalde, Borjan Geshkovski, Domènec Ruiz-Balet

We study the hardmax limit of self-attention dynamics for token embeddings obtained in the zero-temperature () regime, and relate it to the finite- setting. In th…