8 papers
PathWISE: Multi-Agent Cancer Pathway Triaging Ontology Learning from Clinical Flowcharts
Sofiat Abioye, Ufaq Khan, Shazad Ashraf +6
Clinical pathways are disseminated as visual flowcharts where spatial topology, arrow direction, colour coding, and font weight encode critical triage logic that remains inaccessib…
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
Mahdi Alkaeed, Adnan Qayyum, Nabeel Abo Kashreef +2
Large Language Models (LLMs) are increasingly used in clinical applications. However, their behavior remains highly sensitive to subtle linguistic variations, such as rephrasing or…
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
Ezieddin Elmahjub, Junaid Qadir, Abdullah Mushtaq +3
As millions of Muslims turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question arises: Can these AI systems reliably reason about Islamic law? We in…
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
Abdullah Mushtaq, Rafay Naeem, Ezieddin Elmahjub +5
Large language models are increasingly used for Islamic guidance, but risk misquoting texts, misapplying jurisprudence, or producing culturally inconsistent responses. We pilot an…
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
Abdullah Mushtaq, Muhammad Rafay Naeem, Ibrahim Ghaznavi +3
Systematic Literature Reviews (SLRs) are foundational to evidence-based research but remain labor-intensive and prone to inconsistency across disciplines. We present an LLM-based S…
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
Abdullah Mushtaq, Imran Taj, Rafay Naeem +2
Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenizatio…