1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Xinghao Wang, Pengyu Wang, Xiaoran Liu +4
Block-sparse attention is promising for accelerating long-context LLM pre-filling, yet identifying relevant blocks efficiently remains a bottleneck. Existing methods typically empl…