6 papers
Bulbul: A Dataset for Dialectal Arabic Speech Recognition
Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas +30
Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets oft…
LeGo-Code: Can Modular Curriculum Learning Advance Complex Code Generation? Insights from Text-to-SQL
Salmane Chafik, Saad Ezzini, Ismail Berrada
Recently, code-oriented large language models (LLMs) have demonstrated strong capabilities in translating natural language into executable code. Text-to-SQL is a significant applic…
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
Ahmed Hasanaath, Aisha Alansari, Ahmed Ashraf +3
Large language models (LLMs) have shown remarkable progress in reasoning abilities and general natural language processing (NLP) tasks, yet their performance on Arabic data, charac…
AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects
Maram Alharbi, Salmane Chafik, Saad Ezzini +3
The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address thi…
M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text
Salima Lamsiyah, Saad Ezzini, Abdelkader El Mahdaouy +5
The generation of highly fluent text by Large Language Models (LLMs) poses a significant challenge to information integrity and academic research. In this paper, we introduce the M…
Dialect2SQL: A Novel Text-to-SQL Dataset for Arabic Dialects with a Focus on Moroccan Darija
Salmane Chafik, Saad Ezzini, Ismail Berrada
The task of converting natural language questions (NLQs) into executable SQL queries, known as text-to-SQL, has gained significant interest in recent years, as it enables non-techn…