8 citations · 19 across the 3 of their papers we have counts for
4 papers
Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga +11
There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing…
JABER and SABER: Junior and Senior Arabic BERt
Abbas Ghaddar, Yimeng Wu, Ahmad Rashid +10
Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that prev…
ALP-KD: Attention-Based Layer Projection for Knowledge Distillation
Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh +1
Knowledge distillation is considered as a training and compression strategy in which two neural networks, namely a teacher and a student, are coupled together during training. The…
Why Skip If You Can Combine: A Simple Knowledge Distillation Technique for Intermediate Layers
Yimeng Wu, Peyman Passban, Mehdi Rezagholizade +1
With the growth of computing power neural machine translation (NMT) models also grow accordingly and become better. However, they also become harder to deploy on edge devices due t…