collaborators

7 papers

cs.CL2026

CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic Speech

Youssef Saidi, Haroun Elleuch, Fethi Bougares

End-to-end speech Named Entity Recognition (NER) aims to directly extract entities from speech. Prior work has shown that end-to-end (E2E) approaches can outperform cascaded pipeli…

cs.CL2026

SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding

Haroun Elleuch, Salima Mdhaffar, Yannick Estève +1

Spoken Language Understanding (SLU) aims to extract the semantic information from the speech utterance of user queries. It is a core component in a task-oriented dialogue system. W…

cs.CL2026

Ara-Best-RQ: Multi Dialectal Arabic SSL

Haroun Elleuch, Ryan Whetten, Salima Mdhaffar +2

We present Ara-BEST-RQ, a family of self-supervised learning (SSL) models specifically designed for multi-dialectal Arabic speech processing. Leveraging 5,640 hours of crawled Crea…

cs.CL2025

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

Fethi Bougares, Salima Mdhaffar, Haroun Elleuch +1

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the…

cs.CL2025

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

Haroun Elleuch, Youssef Saidi, Salima Mdhaffar +2

This paper describes Elyadata \& LIA's joint submission to the NADI multi-dialectal Arabic Speech Processing 2025. We participated in the Spoken Arabic Dialect Identification (ADI)…

cs.CL2025

ADI-20: Arabic Dialect Identification dataset and models

Haroun Elleuch, Salima Mdhaffar, Yannick Estève +1

We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises…