7 papers
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
Ezieddin Elmahjub, Junaid Qadir, Abdullah Mushtaq +3
As millions of Muslims turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question arises: Can these AI systems reliably reason about Islamic law? We in…
Can AI Chatbots Provide Coaching in Engineering? Beyond Information Processing Toward Mastery
Junaid Qadir, Muhammad Adil Attique, Saleha Shoaib +1
Engineering education faces a double disruption: traditional apprenticeship models that cultivated judgment and tacit skill are eroding, just as generative AI emerges as an informa…
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
Abdullah Mushtaq, Rafay Naeem, Ezieddin Elmahjub +5
Large language models are increasingly used for Islamic guidance, but risk misquoting texts, misapplying jurisprudence, or producing culturally inconsistent responses. We pilot an…
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
Abdullah Mushtaq, Muhammad Rafay Naeem, Ibrahim Ghaznavi +3
Systematic Literature Reviews (SLRs) are foundational to evidence-based research but remain labor-intensive and prone to inconsistency across disciplines. We present an LLM-based S…
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
Abdullah Mushtaq, Imran Taj, Rafay Naeem +2
Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenizatio…
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
Abdullah Mushtaq, Muhammad Rafay Naeem, Muhammad Imran Taj +2
As large language models (LLMs) like GPT-4 and Llama 3 become integral to educational contexts, concerns are mounting over the cultural biases, power imbalances, and ethical limita…