activity
20182022
most citedEnd-to-End Self-Debiasing Framework for Robust NLU Training

29 citations · 65 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CL20226 cited

Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding

Abbas Ghaddar, Yimeng Wu, Sunyam Bagga +11

There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing…

cs.CL20223 cited

CILDA: Contrastive Data Augmentation using Intermediate Layer Knowledge Distillation

Md Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar +3

Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by le…

cs.CL20225 cited

JABER and SABER: Junior and Senior Arabic BERt

Abbas Ghaddar, Yimeng Wu, Ahmad Rashid +10

Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that prev…

cs.CL20211 cited

RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation

Md Akmal Haidar, Nithin Anchuri, Mehdi Rezagholizadeh +3

Intermediate layer knowledge distillation (KD) can improve the standard KD technique (which only targets the output of teacher and student models) especially over large pre-trained…

cs.CL2021

Knowledge Distillation with Noisy Labels for Natural Language Understanding

Shivendra Bhardwaj, Abbas Ghaddar, Ahmad Rashid +5

Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of…

cs.CL202129 cited

End-to-End Self-Debiasing Framework for Robust NLU Training

Abbas Ghaddar, Philippe Langlais, Mehdi Rezagholizadeh +1

Existing Natural Language Understanding (NLU) models have been shown to incorporate dataset biases leading to strong performance on in-distribution (ID) test sets but poor performa…