1 paper · 1 filter
Rakshith Jayanth, Viktor Prasanna
In long-context large language model (LLM) inference, the prefill stage dominates computation due to self-attention over the complete input context. Sparse attention significantly…