23 citations · 36 across the 12 of their papers we have counts for
8 papers · 1 filter
Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory
Sam Buchanan, Druv Pai, Peng Wang +1
In the current era of deep learning and especially generative models, there is significant investment in training very large deep neural networks. Thus far, such models have been "…
On the Edge of Memorization in Diffusion Models
Sam Buchanan, Druv Pai, Yi Ma +1
When do diffusion models reproduce their training data, and when are they able to generate samples beyond it? A practically relevant theoretical understanding of this interplay bet…
Attention-Only Transformers via Unrolled Subspace Denoising
Peng Wang, Yifu Lu, Yaodong Yu +3
Despite the popularity of transformers in practice, their architectures are empirically designed and neither mathematically justified nor interpretable. Moreover, as indicated by m…
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Ziyang Wu, Tianjiao Ding, Yifu Lu +6
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
Tianyu Guo, Druv Pai, Yu Bai +3
Practitioners have consistently observed three puzzling phenomena in transformer-based large language models (LLMs): attention sinks, value-state drains, and residual-state peaks,…
A Global Geometric Analysis of Maximal Coding Rate Reduction
Peng Wang, Huikang Liu, Druv Pai +4
The maximal coding rate reduction (MCR) objective for learning structured and compact deep representations is drawing increasing attention, especially after its recent usage in…