Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
Zhengfu He, Junxuan Wang, Rui Lin +5
We propose Low-Rank Sparse Attention (Lorsa), a sparse replacement model of Transformer attention layers to disentangle original Multi Head Self Attention (MHSA) into individually…
cs.LG2023
Lite it fly: An All-Deformable-Butterfly Network
Rui Lin, Jason Chun Lok Li, Jiajun Zhou +3
Most deep neural networks (DNNs) consist fundamentally of convolutional and/or fully connected layers, wherein the linear transform can be cast as the product between a filter matr…