activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas +30

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets oft…

cs.CL2026

LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

Omar El Bachyr, Fred Philippy, Laura Maria Bernardy +3

Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-lingual, low-resource settings…

cs.CL2026

Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection

Rasha Albalawi, Nuha Albadi, Hamzah Luqman +4

Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-XT,…

cs.CL2026

Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG

Omar El Bachyr, Yewei Song, Saad Ezzini +5

PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, tables, and images, poses signifi…

cs.CL2025

AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP

Ahmed Hasanaath, Aisha Alansari, Ahmed Ashraf +3

Large language models (LLMs) have shown remarkable progress in reasoning abilities and general natural language processing (NLP) tasks, yet their performance on Arabic data, charac…

cs.CL2025

AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects

Maram Alharbi, Salmane Chafik, Saad Ezzini +3

The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address thi…