12 citations · 39 across the 11 of their papers we have counts for
10 papers
Multimodal Audio-textual Architecture for Robust Spoken Language Understanding
Anderson R. Avila, Mehdi Rezagholizadeh, Chao Xing
Recent voice assistants are usually based on the cascade spoken language understanding (SLU) solution, which consists of an automatic speech recognition (ASR) engine and a natural…
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo +6
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval…
Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga +11
There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing…
CILDA: Contrastive Data Augmentation using Intermediate Layer Knowledge Distillation
Md Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar +3
Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by le…
When Chosen Wisely, More Data Is What You Need: A Universal Sample-Efficient Strategy For Data Augmentation
Ehsan Kamalloo, Mehdi Rezagholizadeh, Ali Ghodsi
Data Augmentation (DA) is known to improve the generalizability of deep neural networks. Most existing DA techniques naively add a certain number of augmented samples without consi…
JABER and SABER: Junior and Senior Arabic BERt
Abbas Ghaddar, Yimeng Wu, Ahmad Rashid +10
Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that prev…