activity
20202026
most citedALP-KD: Attention-Based Layer Projection for Knowledge Distillation

8 citations · 19 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2023

AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing

Asaad Alghamdi, Xinyu Duan, Wei Jiang +9

Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). In this work, we pr…

cs.CL20226 cited

Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding

Abbas Ghaddar, Yimeng Wu, Sunyam Bagga +11

There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing…

cs.CL20225 cited

JABER and SABER: Junior and Senior Arabic BERt

Abbas Ghaddar, Yimeng Wu, Ahmad Rashid +10

Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that prev…

cs.CL20208 cited

ALP-KD: Attention-Based Layer Projection for Knowledge Distillation

Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh +1

Knowledge distillation is considered as a training and compression strategy in which two neural networks, namely a teacher and a student, are coupled together during training. The…

cs.CL2020

Why Skip If You Can Combine: A Simple Knowledge Distillation Technique for Intermediate Layers

Yimeng Wu, Peyman Passban, Mehdi Rezagholizade +1

With the growth of computing power neural machine translation (NMT) models also grow accordingly and become better. However, they also become harder to deploy on edge devices due t…