10 papers
Perceptrons and localization of attention's mean-field landscape
Antonio Ãlvarez-López, Borjan Geshkovski, Domènec Ruiz-Balet
The forward pass of a Transformer can be seen as an interacting particle system on the unit sphere: time plays the role of layers, particles that of token embeddings, and the unit…
Constructive conditional normalizing flows
Borjan Geshkovski, Domènec Ruiz-Balet
Motivated by applications in conditional sampling, given a probability measure and a diffeomorphism , we consider the problem of simultaneously approximating and the…
Kinetic theory for Transformers and the lost-in-the-middle phenomenon
Mitia Duerinckx, Borjan Geshkovski, Stefano Rossi
We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle system. Adapting cumulant expansion…
Homogenized Transformers
Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet
We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth…
Measure-to-measure interpolation using Transformers
Borjan Geshkovski, Philippe Rigollet, Domènec Ruiz-Balet
Transformers are deep neural network architectures that underpin the recent successes of large language models. Unlike more classical architectures that can be viewed as point-to-p…
On the number of modes of Gaussian kernel density estimators
Borjan Geshkovski, Philippe Rigollet, Yihang Sun
We consider the Gaussian kernel density estimator with bandwidth $β^{-\frac12}$ of iid Gaussian samples. Using the Kac-Rice formula and an Edgeworth expansion, we prove that t…