most citedADI-20: Arabic Dialect Identification dataset and models

3 citations · 3 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2025

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

Fethi Bougares, Salima Mdhaffar, Haroun Elleuch +1

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the…

cs.CL2025

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

Haroun Elleuch, Youssef Saidi, Salima Mdhaffar +2

This paper describes Elyadata \& LIA's joint submission to the NADI multi-dialectal Arabic Speech Processing 2025. We participated in the Spoken Arabic Dialect Identification (ADI)…

cs.CL20253 cited

ADI-20: Arabic Dialect Identification dataset and models

Haroun Elleuch, Salima Mdhaffar, Yannick Estève +1

We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises…

cs.CL2025

In-domain SSL pre-training and streaming ASR

Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière +6

In this study, we investigate the benefits of domain-specific self-supervised pre-training for both offline and streaming ASR in Air Traffic Control (ATC) environments. We train BE…

cs.CL2025

SENSE models: an open source solution for multilingual and multimodal semantic-based tasks

Salima Mdhaffar, Haroun Elleuch, Chaimae Chellaf +2

This paper introduces SENSE (Shared Embedding for N-lingual Speech and tExt), an open-source solution inspired by the SAMU-XLSR framework and conceptually similar to Meta AI's SONA…