3 papers
cs.LG2025
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
Junxuan Wang, Xuyang Ge, Wentao Shu +2
Transformer architectures, and their attention mechanisms in particular, form the foundation of modern large language models. While transformer models are widely believed to operat…
cs.LG2025
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
Zhengfu He, Junxuan Wang, Rui Lin +5
We propose Low-Rank Sparse Attention (Lorsa), a sparse replacement model of Transformer attention layers to disentangle original Multi Head Self Attention (MHSA) into individually…
cs.CL2024
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
Junxuan Wang, Xuyang Ge, Wentao Shu +4
The hypothesis of Universality in interpretability suggests that different neural networks may converge to implement similar algorithms on similar tasks. In this work, we investiga…