65 citations · 65 across the 1 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023
TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer
Zhen Qin, Dong Li, Weigao Sun +8
We present TransNormerLLM, the first linear attention-based Large Language Model (LLM) that outperforms conventional softmax attention-based models in terms of both accuracy and ef…
cs.CL2022★ 65 cited
cosFormer: Rethinking Softmax in Attention
Zhen Qin, Weixuan Sun, Hui Deng +6
Transformer has shown great successes in natural language processing, computer vision, and audio processing. As one of its core components, the softmax attention helps to capture l…