22 citations · 34 across the 7 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022★ 1 cited
Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation
Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li +1
As its core computation, a self-attention mechanism gauges pairwise correlations across the entire input sequence. Despite favorable performance, calculating pairwise correlations…
cs.CL2022
Accelerating Attention through Gradient-Based Learned Runtime Pruning
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh +2
Self-attention is a key enabler of state-of-art accuracy for various transformer-based Natural Language Processing models. This attention mechanism calculates a correlation score f…