collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling

Tajwaar Shafiq, Hunzalah Hassan Bhatti, Firoj Alam +1

We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate four omni LLMs in a zero-shot setting an…

cs.CL2026

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti +11

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question ans…

cs.CL2026

Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models

Basel Mousi, Fahim Dalvi, Shammur Chowdhury +2

Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whe…

cs.CL2026

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti +2

Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses,…

cs.CL2026

Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation

Firoj Alam, Gagan Bhatia, Sahinur Rahman Laskar +1

While Large Language Models (LLMs) are increasingly adopted as automated judges for evaluating generated text, their outputs are often costly, and highly sensitive to prompt design…

cs.CL2026

Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models

Basel Mousi, Fahim Dalvi, Shammur Chowdhury +2

Vision-language models (VLMs) can achieve high accuracy while still accepting culturally plausible but visually incorrect interpretations. Existing hallucination benchmarks rarely…