1 citations · 1 across the 9 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
Qihang Fan, Huaibo Huang, Zhiying Wu +2
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-in…
cs.CL2026
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
Qihang Fan, Huaibo Huang, Zhiying Wu +3
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-in…