10 citations · 10 across the 1 of their papers we have counts for
3 papers
KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation
Marzieh S. Tahaei, Ella Charlaix, Vahid Partovi Nia +2
The development of over-parameterized pre-trained language models has made a significant contribution toward the success of natural language processing. While over-parameterization…
Block Pruning For Faster Transformers
François Lagunas, Ella Charlaix, Victor Sanh +1
Pre-training has improved model accuracy for both classification and generation tasks at the cost of introducing much larger and slower models. Pruning methods have proven to be an…
Fully Quantized Transformer for Machine Translation
Gabriele Prato, Ella Charlaix, Mehdi Rezagholizadeh
State-of-the-art neural machine translation methods employ massive amounts of parameters. Drastically reducing computational costs of such methods without affecting performance has…