9 papers
Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework
Seyed Amir Ahmad Safavi-Naini, Elahe Meftah, Josh Mohess +11
The competency of any intelligent agent is bounded by its formal account of the world in which it operates. Clinical AI lacks such an account. Existing frameworks address evaluatio…
Kantian-Utilitarian XAI: Meta-Explained
Zahra Atf, Peter R. Lewis
We present a gamified explainable AI (XAI) system for ethically aware consumer decision-making in the coffee domain. Each session comprises six rounds with three options per round.…
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
Zahra Atf, Peter R Lewis
ScenarioBench is a policy-grounded, trace-aware benchmark for evaluating Text-to-SQL and retrieval-augmented generation in compliance contexts. Each YAML scenario includes a no-pee…
Large Language Models versus Classical Machine Learning: Performance in COVID-19 Mortality Prediction Using High-Dimensional Tabular Data
Mohammadreza Ghaffarzadeh-Esfahani, Mahdi Ghaffarzadeh-Esfahani, Arian Salahi-Niri +39
This study compared the performance of classical feature-based machine learning models (CMLs) and large language models (LLMs) in predicting COVID-19 mortality using high-dimension…
Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
Zahra Atf, Peter R Lewis
Large language models (LLMs) are increasingly used in high-stakes settings, where explaining uncertainty is both technical and ethical. Probabilistic methods are often opaque and m…
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
Nariman Naderi, Zahra Atf, Peter R Lewis +3
This paper investigates how prompt engineering techniques impact both accuracy and confidence elicitation in Large Language Models (LLMs) applied to medical contexts. Using a strat…