1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CL2026★ 1 cited
Why Softmax Attention Outperforms Linear Attention
Yichuan Deng, Zhao Song, Kaijun Yuan +1
Large transformer models have achieved state-of-the-art results in numerous natural language processing tasks. Among the pivotal components of the transformer architecture, the att…
cs.LG2025
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally -Sparse
Yichuan Deng, Zhao Song, Jing Xiong +1
Sparse Attention is a technique that approximates standard attention computation with sub-quadratic complexity. This is achieved by selectively ignoring smaller entries in the atte…