1 paper
Hongyin Tang, Di Xiu, Lanrui Wang +3
The quadratic computational complexity of the attention mechanism in current Large Language Models (LLMs) renders inference with long contexts prohibitively expensive. To address t…