14 citations · 24 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022★ 4 cited
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers
Zhewei Yao, Xiaoxia Wu, Conglong Li +4
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP. However, those large models also introduce prohib…
cs.CL2021★ 6 cited
NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMM
Connor Holmes, Minjia Zhang, Yuxiong He +1
Natural Language Processing (NLP) has recently achieved success by using huge pre-trained Transformer networks. However, these models often contain hundreds of millions or even bil…