57 citations · 60 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 57 cited
Efficiently Scaling Transformer Inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery +7
We study the problem of efficient generative inference for Transformer models, in one of its most challenging settings: large deep models, with tight latency targets and long seque…
cs.LG2022★ 3 cited
Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask
Sheng-Chun Kao, Amir Yazdanbakhsh, Suvinay Subramanian +3
Sparsity has become one of the promising methods to compress and accelerate Deep Neural Networks (DNNs). Among different categories of sparsity, structured sparsity has gained more…