11 papers
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
Samet Demir, Zafer Dogan
Pretrained Transformers can perform in-context learning (ICL) from a few demonstrations, but this ability can fail sharply when the test distribution differs from pretraining, a co…
Learnability and Competition in High-Dimensional Multi-Component ICA
Eser Ilke Genc, Samet Demir, Zafer Dogan
Independent Component Analysis (ICA) is a foundational tool for unsupervised representation learning, yet its high-dimensional theory remains largely limited to single-component re…
Input-Label Correlation Governs a Linear-to-Nonlinear Transition in Random Features under Spiked Covariance
Samet Demir, Zafer Dogan
Random feature models (RFMs), two-layer networks with a randomly initialized fixed first layer and a trained linear readout, are among the simplest nonlinear predictors. Prior asym…
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
Samet Demir, Zafer Dogan
Pretrained Transformers demonstrate remarkable in-context learning (ICL) capabilities, enabling them to adapt to new tasks from demonstrations without parameter updates. However, t…
Learning Beyond the Gaussian Data: Learning Dynamics of Neural Networks on an Expressive and Cumulant-Controllable Data Model
Onat Ure, Samet Demir, Zafer Dogan
We study the effect of high-order statistics of data on the learning dynamics of neural networks (NNs) by using a moment-controllable non-Gaussian data model. Considering the expre…
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
Samet Demir, Zafer Dogan
Many online learning algorithms, including classical online PCA methods, enforce explicit normalization steps that discard the evolving norm of the parameter vector. We show that t…