activity
20242026
collaborators

8 papers

stat.ML2026

Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch

Mary Letey, Yue M. Lu, Cengiz Pehlevan +1

Modern sequence models have a striking capacity for in-context learning (ICL); they can perform new tasks based only on examples given in the prompt. Understanding how this ability…

stat.ML2026

Sharp Capacity Thresholds in Linear Associative Memory: From Top-1 Retrieval to Tail-Average Learning

Nicholas Barnfield, Juno Kim, Eshaan Nichani +2

How many key-value associations can a linear memory store? The answer depends not only on the degrees of freedom in the memory matrix, but also on the retrieval c…

stat.ML2026

Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active Learning

Hugo Cui, Yue M. Lu

We study a class of iterated empirical risk minimization (ERM) procedures in which two successive ERMs are performed on the same dataset, and the predictions of the first estimator…

cs.LG2025

A solvable model of learning generative diffusion: theory and insights

Hugo Cui, Cengiz Pehlevan, Yue M. Lu

In this manuscript, we consider the problem of learning a flow or diffusion-based generative model parametrized by a two-layer auto-encoder, trained with online stochastic gradient…

stat.ML2025

Asymptotic theory of in-context learning by linear attention

Yue M. Lu, Mary I. Letey, Jacob A. Zavatone-Veth +2

Transformers have a remarkable ability to learn and execute tasks based on examples provided within the input itself, without explicit prior training. It has been argued that this…

stat.ML2025

Pretrain-Test Task Alignment Governs Generalization in In-Context Learning

Mary I. Letey, Jacob A. Zavatone-Veth, Yue M. Lu +1

In-context learning (ICL) is a central capability of Transformer models, but the structures in data that enable its emergence and govern its robustness remain poorly understood. In…