1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Siran Liu, Yang Xue, Theo Tang +11
Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-K stage must still process materialized score r…