1 paper
Dan Peng, Zhihui Fu, Zewen Ye +2
Sparse attention methods exploit the inherent sparsity in attention to speed up the prefilling phase of long-context inference, mitigating the quadratic complexity of full attentio…