collaborators

15 papers

cs.LG2026

Measure-to-measure Regression with Transformers

Matthew Vandergrift, Martha White, Yury Polyanskiy +2

Many learning problems require predicting how populations evolve under an unknown transformation. A natural representation for such populations is a probability measure, with point…

cs.LG2026

Propagation of Chaos in Contextual Flow Maps

Shi Chen, Zhengjiang Lin, Kaizhao Liu +1

We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a…

cs.LG2026

Scaling Limits of Long-Context Transformers

Giuseppe Bruno, Shi Chen, Zhengjiang Lin +2

We study the long-context limit of softmax self-attention with a fixed query and a random context of i.i.d. keys on the sphere, viewing the inverse temperature as the sc…

cs.LG2026

Quantitative Clustering in Mean-Field Transformer Models

Shi Chen, Zhengjiang Lin, Yury Polyanskiy +1

The evolution of tokens through deep transformer models can be modeled as an interacting particle system that has been shown to exhibit an asymptotic clustering behavior akin to th…

math.PR2026

Homogenized Transformers

Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet

We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth…

cs.LG2026

YuriiFormer: A Suite of Nesterov-Accelerated Transformers

Aleksandr Zimin, Yury Polyanskiy, Philippe Rigollet

We propose a variational framework that interprets transformer layers as iterations of an optimization algorithm acting on token embeddings. In this view, self-attention implements…