collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs

Mahdi Alkaeed, Adnan Qayyum, Nabeel Abo Kashreef +2

Large Language Models (LLMs) are increasingly used in clinical applications. However, their behavior remains highly sensitive to subtle linguistic variations, such as rephrasing or…

cs.CL2026

IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions

Ezieddin Elmahjub, Junaid Qadir, Abdullah Mushtaq +3

As millions of Muslims turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question arises: Can these AI systems reliably reason about Islamic law? We in…

cs.CL2025

Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content

Abdullah Mushtaq, Rafay Naeem, Ezieddin Elmahjub +5

Large language models are increasingly used for Islamic guidance, but risk misquoting texts, misapplying jurisprudence, or producing culturally inconsistent responses. We pilot an…

cs.CL2025

WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models

Abdullah Mushtaq, Imran Taj, Rafay Naeem +2

Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenizatio…

cs.CL2025

Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens

Abdullah Mushtaq, Muhammad Rafay Naeem, Muhammad Imran Taj +2

As large language models (LLMs) like GPT-4 and Llama 3 become integral to educational contexts, concerns are mounting over the cultural biases, power imbalances, and ethical limita…