1 paper
Siyu Li, Dong Wang, Jie Zhou +3
Post-hoc sparse attention accelerates long-context prefill by routing each query to a small set of token-level interactions. Hard selection, however, assigns zero probability to ev…