1 citations · 1 across the 3 of their papers we have counts for
5 papers
The Effect of Training Task Diversity on In-Context Learning through the Lens of Low-Dimensional Subspaces
Soo Min Kwon, Alec S. Xu, Can Yaras +3
The transformer's emergent ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its underlying mechanisms. Existing works often s…
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
Soo Min Kwon, Alec S. Xu, Can Yaras +2
The transformer's remarkable ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its strengths and limitations. However, a theor…
Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension
Alec S. Xu, Can Yaras, Peng Wang +1
Deep neural networks have attained remarkable success across diverse classification tasks. Recent empirical studies have shown that deep networks learn features that are linearly s…
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
Alec S. Xu, Can Yaras, Matthew Asato +2
Recent empirical evidence has demonstrated that the training dynamics of large-scale deep neural networks occur within low-dimensional subspaces. While this has inspired new resear…
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
Can Yaras, Alec S. Xu, Pierre Abillama +2
Transformers have achieved state-of-the-art performance across various tasks, but suffer from a notable quadratic complexity in sequence length due to the attention mechanism. In t…