4 citations · 5 across the 7 of their papers we have counts for
1 paper · 2 filters
Siran Liu, Yang Xue, Theo Tang +11
Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-K stage must still process materialized score r…