1 paper · 1 filter
Sirui Chen, Jingji Chen, Siqi Zhu +3
Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited parallelism or incur high communication costs…