Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification
Kimia Hamidieh, Veronika Thost, Walter Gerych +2
Large language models (LLMs) often produce confident yet incorrect responses, and uncertainty quantification is one potential solution to more robust usage. Recent works routinely…
cs.AI2025
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
Abinitha Gourabathina, Yuexing Hao, Walter Gerych +1
Clinical robustness is critical to the safe deployment of medical Large Language Models (LLMs), but key questions remain about how LLMs and humans may differ in response to the rea…