activity
20192024
most citedMaking a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

12 citations · 39 across the 11 of their papers we have counts for

collaborators

10 papers

cs.CL20231 cited

Multimodal Audio-textual Architecture for Robust Spoken Language Understanding

Anderson R. Avila, Mehdi Rezagholizadeh, Chao Xing

Recent voice assistants are usually based on the cascade spoken language understanding (SLU) solution, which consists of an automatic speech recognition (ASR) engine and a natural…

cs.IR202212 cited

Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo +6

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval…

cs.CL20226 cited

Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding

Abbas Ghaddar, Yimeng Wu, Sunyam Bagga +11

There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing…

cs.CL20223 cited

CILDA: Contrastive Data Augmentation using Intermediate Layer Knowledge Distillation

Md Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar +3

Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by le…

cs.LG2022

When Chosen Wisely, More Data Is What You Need: A Universal Sample-Efficient Strategy For Data Augmentation

Ehsan Kamalloo, Mehdi Rezagholizadeh, Ali Ghodsi

Data Augmentation (DA) is known to improve the generalizability of deep neural networks. Most existing DA techniques naively add a certain number of augmented samples without consi…

cs.CL20225 cited

JABER and SABER: Junior and Senior Arabic BERt

Abbas Ghaddar, Yimeng Wu, Ahmad Rashid +10

Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that prev…