1 paper
Xiaowei Ye, Xiaoyu He, Chao Liao +2
Transformers serve as the foundation of most modern large language models. To mitigate the quadratic complexity of standard full attention, various efficient attention mechanisms,…