1 paper
Yixing Xu, Shivank Nag, Dong Li +2
Transformer-based LLMs have achieved exceptional performance across a wide range of NLP tasks. However, the standard self-attention mechanism suffers from quadratic time complexity…