15 papers
Measure-to-measure Regression with Transformers
Matthew Vandergrift, Martha White, Yury Polyanskiy +2
Many learning problems require predicting how populations evolve under an unknown transformation. A natural representation for such populations is a probability measure, with point…
Propagation of Chaos in Contextual Flow Maps
Shi Chen, Zhengjiang Lin, Kaizhao Liu +1
We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a…
Scaling Limits of Long-Context Transformers
Giuseppe Bruno, Shi Chen, Zhengjiang Lin +2
We study the long-context limit of softmax self-attention with a fixed query and a random context of i.i.d. keys on the sphere, viewing the inverse temperature as the sc…
Quantitative Clustering in Mean-Field Transformer Models
Shi Chen, Zhengjiang Lin, Yury Polyanskiy +1
The evolution of tokens through deep transformer models can be modeled as an interacting particle system that has been shown to exhibit an asymptotic clustering behavior akin to th…
Homogenized Transformers
Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet
We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth…
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
Aleksandr Zimin, Yury Polyanskiy, Philippe Rigollet
We propose a variational framework that interprets transformer layers as iterations of an optimization algorithm acting on token embeddings. In this view, self-attention implements…