1 paper
Bin Li, Sisi Liu, Chenyang Hu +3
Dynamic sparse attention reduces long-context prefill cost by routing each query chunk to a small set of key chunks at every Transformer layer. The sparse attention kernel avoids m…