5 papers · 1 filter
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
Kareem Elozeiri, Mervat Abassy, Omar Kallas +4
A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLMs to overproduce MSA and struggle with dialectally accurate gen…
Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
Yuxia Wang, Rui Xing, Jonibek Mansurov +23
Prior studies have shown that distinguishing text generated by Large Language Models (LLMs) from human-written one is highly challenging for humans, and often no better than random…
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
Kareem Elozeiri, Mervat Abassy, Preslav Nakov +1
Commonsense validation evaluates whether a sentence aligns with everyday human understanding, a critical capability for developing robust natural language understanding systems. Wh…
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection
Mervat Abassy, Kareem Elozeiri, Alexander Aziz +21
The ease of access to large language models (LLMs) has enabled a widespread of machine-generated texts, and now it is often hard to tell whether a piece of text was human-written o…
GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human
Yuxia Wang, Artem Shelmanov, Jonibek Mansurov +23
We present the GenAI Content Detection Task~1 -- a shared task on binary machine generated text detection, conducted as a part of the GenAI workshop at COLING 2025. The task consis…