activity
20122022
most citedA Berkeley View of Systems Challenges for AI

176 citations · 547 across the 34 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL20212 cited

Pro-KD: Progressive Distillation by Following the Footsteps of the Teacher

Mehdi Rezagholizadeh, Aref Jafari, Puneeth Salad +3

With ever growing scale of neural models, knowledge distillation (KD) attracts more attention as a prominent tool for neural model compression. However, there are counter intuitive…

cs.CL2021

Knowledge Distillation with Noisy Labels for Natural Language Understanding

Shivendra Bhardwaj, Abbas Ghaddar, Ahmad Rashid +5

Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of…

cs.CL20213 cited

How to Select One Among All? An Extensive Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding

Tianda Li, Ahmad Rashid, Aref Jafari +3

Knowledge Distillation (KD) is a model compression algorithm that helps transfer the knowledge of a large neural network into a smaller one. Even though KD has shown promise on a w…

cs.CL202110 cited

KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation

Marzieh S. Tahaei, Ella Charlaix, Vahid Partovi Nia +2

The development of over-parameterized pre-trained language models has made a significant contribution toward the success of natural language processing. While over-parameterization…

cs.CL2021

Annealing Knowledge Distillation

Aref Jafari, Mehdi Rezagholizadeh, Pranav Sharma +1

Significant memory and computational requirements of large deep neural networks restrict their application on edge devices. Knowledge distillation (KD) is a prominent model compres…

cs.CL2020

Segmentation Approach for Coreference Resolution Task

Aref Jafari, Ali Ghodsi

In coreference resolution, it is important to consider all members of a coreference cluster and decide about all of them at once. This technique can help to avoid losing precision…