1 paper · 1 filter
Siran Liu, Yang Xue, Theo Tang +11
Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-K stage must still process materialized score r…