30 citations · 64 across the 6 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.CL2022
Compressing Pre-trained Transformers via Low-Bit NxM Sparsity for Natural Language Understanding
Connor Holmes, Minjia Zhang, Yuxiong He +1
In recent years, large pre-trained Transformer networks have demonstrated dramatic improvements in many natural language understanding tasks. However, the huge size of these models…
cs.LG2022★ 2 cited
ScaLA: Accelerating Adaptation of Pre-Trained Transformer-Based Language Models via Efficient Large-Batch Adversarial Noise
Minjia Zhang, Niranjan Uma Naresh, Yuxiong He
In recent years, large pre-trained Transformer-based language models have led to dramatic improvements in many natural language understanding tasks. To train these models with incr…