1 paper
Ziming Wang, Xiang Wang, Kailong Peng +5
Large Language Models (LLMs) encounter significant performance bottlenecks in long-sequence tasks due to the computational complexity and memory overhead inherent in the self-atten…