4 citations · 4 across the 1 of their papers we have counts for
1 paper
Rajiv Movva, Jinhao Lei, Shayne Longpre +2
Quantization, knowledge distillation, and magnitude pruning are among the most popular methods for neural network compression in NLP. Independently, these methods reduce model size…