2 citations · 4 across the 7 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
Zahra Atf, Peter R Lewis
ScenarioBench is a policy-grounded, trace-aware benchmark for evaluating Text-to-SQL and retrieval-augmented generation in compliance contexts. Each YAML scenario includes a no-pee…
cs.CL2025
Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
Zahra Atf, Peter R Lewis
Large language models (LLMs) are increasingly used in high-stakes settings, where explaining uncertainty is both technical and ethical. Probabilistic methods are often opaque and m…
cs.CL2025
Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
Nariman Naderi, Seyed Amir Ahmad Safavi-Naini, Thomas Savage +4
This study evaluated self-reported response certainty across several large language models (GPT, Claude, Llama, Phi, Mistral, Gemini, Gemma, and Qwen) using 300 gastroenterology bo…