activity
20192026
most citedOn neural network kernels and the storage capacity problem

8 citations · 10 across the 17 of their papers we have counts for

collaborators
Showing stat.MLShow all

15 papers · 1 filter

stat.ML2026

Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch

Mary Letey, Yue M. Lu, Cengiz Pehlevan +1

Modern sequence models have a striking capacity for in-context learning (ICL); they can perform new tasks based only on examples given in the prompt. Understanding how this ability…

stat.ML2026

An Asymptotic Theory of Chain-of-Thought in In-Context Learning

Kaito Takanami, Cengiz Pehlevan

Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermediate reasoning steps at infere…

stat.ML2025

Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time

Blake Bordelon, Mary I. Letey, Cengiz Pehlevan

We study in-context learning (ICL) of linear regression in a deep linear self-attention model, characterizing how performance depends on various computational and statistical resou…

stat.ML2025

Pretrain-Test Task Alignment Governs Generalization in In-Context Learning

Mary I. Letey, Jacob A. Zavatone-Veth, Yue M. Lu +1

In-context learning (ICL) is a central capability of Transformer models, but the structures in data that enable its emergence and govern its robustness remain poorly understood. In…

stat.ML2024

How Feature Learning Can Improve Neural Scaling Laws

Blake Bordelon, Alexander Atanasov, Cengiz Pehlevan

We develop a solvable model of neural scaling laws beyond the kernel limit. Theoretical analysis of this model shows how performance scales with model size, training time, and the…

stat.ML2024

Risk and cross validation in ridge regression with correlated samples

Alexander Atanasov, Jacob A. Zavatone-Veth, Cengiz Pehlevan

Recent years have seen substantial advances in our understanding of high-dimensional ridge regression, but existing theories assume that training examples are independent. By lever…