activity
20232026
most citedWhite-Box Transformers via Sparse Rate Reduction

23 citations · 36 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory

Sam Buchanan, Druv Pai, Peng Wang +1

In the current era of deep learning and especially generative models, there is significant investment in training very large deep neural networks. Thus far, such models have been "…

cs.LG2025

On the Edge of Memorization in Diffusion Models

Sam Buchanan, Druv Pai, Yi Ma +1

When do diffusion models reproduce their training data, and when are they able to generate samples beyond it? A practically relevant theoretical understanding of this interplay bet…

cs.LG2025

Attention-Only Transformers via Unrolled Subspace Denoising

Peng Wang, Yifu Lu, Yaodong Yu +3

Despite the popularity of transformers in practice, their architectures are empirically designed and neither mathematically justified nor interpretable. Moreover, as indicated by m…

cs.LG20247 cited

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

Ziyang Wu, Tianjiao Ding, Yifu Lu +6

The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…

cs.LG2024

Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs

Tianyu Guo, Druv Pai, Yu Bai +3

Practitioners have consistently observed three puzzling phenomena in transformer-based large language models (LLMs): attention sinks, value-state drains, and residual-state peaks,…

cs.LG2024

A Global Geometric Analysis of Maximal Coding Rate Reduction

Peng Wang, Huikang Liu, Druv Pai +4

The maximal coding rate reduction (MCR) objective for learning structured and compact deep representations is drawing increasing attention, especially after its recent usage in…