1 citations · 2 across the 7 of their papers we have counts for
8 papers · 1 filter
Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination
Peng Wang, Xiao Li, Can Yaras +4
Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remains an open question how deep networks…
Stochastic Sparse Attention for Memory-Bound Inference
Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5
Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all key and value vectors from KV cache. We present Stochastic A…
Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension
Alec S. Xu, Can Yaras, Peng Wang +1
Deep neural networks have attained remarkable success across diverse classification tasks. Recent empirical studies have shown that deep networks learn features that are linearly s…
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
Can Yaras, Siyi Chen, Peng Wang +1
Multimodal learning has recently gained significant popularity, demonstrating impressive performance across various zero-shot classification tasks and a range of perceptive and gen…
Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation
Can Yaras, Peng Wang, Laura Balzano +1
While overparameterization in machine learning models offers great benefits in terms of optimization and generalization, it also leads to increased computational requirements as mo…
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
Alec S. Xu, Can Yaras, Matthew Asato +2
Recent empirical evidence has demonstrated that the training dynamics of large-scale deep neural networks occur within low-dimensional subspaces. While this has inspired new resear…