31 citations · 33 across the 5 of their papers we have counts for
5 papers
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
Zahra Atf, Peter R Lewis
ScenarioBench is a policy-grounded, trace-aware benchmark for evaluating Text-to-SQL and retrieval-augmented generation in compliance contexts. Each YAML scenario includes a no-pee…
Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
Zahra Atf, Peter R Lewis
Large language models (LLMs) are increasingly used in high-stakes settings, where explaining uncertainty is both technical and ethical. Probabilistic methods are often opaque and m…
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
Nariman Naderi, Zahra Atf, Peter R Lewis +3
This paper investigates how prompt engineering techniques impact both accuracy and confidence elicitation in Large Language Models (LLMs) applied to medical contexts. Using a strat…
Is Trust Correlated With Explainability in AI? A Meta-Analysis
Zahra Atf, Peter R. Lewis
This study critically examines the commonly held assumption that explicability in artificial intelligence (AI) systems inherently boosts user trust. Utilizing a meta-analytical app…
The challenge of uncertainty quantification of large language models in medicine
Zahra Atf, Seyed Amir Ahmad Safavi-Naini, Peter R. Lewis +4
This study investigates uncertainty quantification in large language models (LLMs) for medical applications, emphasizing both technical innovations and philosophical implications.…