10 papers
Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination
Peng Wang, Xiao Li, Can Yaras +4
Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remains an open question how deep networks…
The Effect of Training Task Diversity on In-Context Learning through the Lens of Low-Dimensional Subspaces
Soo Min Kwon, Alec S. Xu, Can Yaras +3
The transformer's emergent ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its underlying mechanisms. Existing works often s…
Stochastic Sparse Attention for Memory-Bound Inference
Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5
Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all key and value vectors from KV cache. We present Stochastic A…
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
Soo Min Kwon, Alec S. Xu, Can Yaras +2
The transformer's remarkable ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its strengths and limitations. However, a theor…
Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension
Alec S. Xu, Can Yaras, Peng Wang +1
Deep neural networks have attained remarkable success across diverse classification tasks. Recent empirical studies have shown that deep networks learn features that are linearly s…
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
Can Yaras, Siyi Chen, Peng Wang +1
Multimodal learning has recently gained significant popularity, demonstrating impressive performance across various zero-shot classification tasks and a range of perceptive and gen…