9 citations · 30 across the 7 of their papers we have counts for
1 paper · 1 filter
Beidi Chen, Tri Dao, Eric Winsor +3
Recent advances in efficient Transformers have exploited either the sparsity or low-rank properties of attention matrices to reduce the computational and memory bottlenecks of mode…