collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

Mo El-Haj

Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxific…

cs.CL2026

Structured Prompting for Arabic Essay Proficiency: A Trait-Centric Evaluation Approach

Salim Al Mandhari, Hieu Pham Dinh, Mo El-Haj +1

This paper presents a novel prompt engineering framework for trait specific Automatic Essay Scoring (AES) in Arabic, leveraging large language models (LLMs) under zero-shot and few…

cs.CL2026

Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry

Mo El-Haj

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus c…

cs.CL2026

FreeTxt-Vi: A Benchmarked Vietnamese-English Toolkit for Segmentation, Sentiment, and Summarisation

Hung Nguyen Huy, Mo El-Haj, Dawn Knight +1

FreeTxt-Vi is a free and open source web based toolkit for creating and analysing bilingual Vietnamese English text collections. Positioned at the intersection of corpus linguistic…

cs.CL2026

VietJobs: A Vietnamese Job Advertisement Dataset

Hieu Pham Dinh, Hung Nguyen Huy, Mo El-Haj

VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces…

cs.CL2024

AraFinNLP 2024: The First Arabic Financial NLP Shared Task

Sanad Malaysha, Mo El-Haj, Saad Ezzini +5

The expanding financial markets of the Arab world require sophisticated Arabic NLP tools. To address this need within the banking domain, the Arabic Financial NLP (AraFinNLP) share…