7 papers
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
ADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational Sociopragmatics
Hend Al-Khalifa, Nadia Ghezaiel, Maria Bounnit +5
The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse la…
From Code-Centric to Concept-Centric: Teaching NLP with LLM-Assisted "Vibe Coding"
Hend Al-Khalifa
The rapid advancement of Large Language Models (LLMs) presents both challenges and opportunities for Natural Language Processing (NLP) education. This paper introduces ``Vibe Codin…
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
Srikant Panda, Hitesh Laxmichand Patel, Shahad Al-Khalifa +3
Recent evaluations of Large language models (LLMs) audit social bias primarily through prompts that explicitly reference demographic attributes, overlooking whether models infer se…
The Landscape of Arabic Large Language Models (ALLMs): A New Era for Arabic Language Technology
Shahad Al-Khalifa, Nadir Durrani, Hend Al-Khalifa +1
The emergence of ChatGPT marked a transformative milestone for Artificial Intelligence (AI), showcasing the remarkable potential of Large Language Models (LLMs) to generate human-l…
The Prompting Brain: Neurocognitive Markers of Expertise in Guiding Large Language Models
Hend Al-Khalifa, Raneem Almansour, Layan Abdulrahman Alhuasini +4
Prompt engineering has rapidly emerged as a critical skill for effective interaction with large language models (LLMs). However, the cognitive and neural underpinnings of this expe…