4 papers · 1 filter
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Hongjian Zhou, Xinyu Zou, Jinge Wu +19
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…
BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
Sean Wu, Fredrik K. Gustafsson, Edward Phillips +3
Large language models (LLMs) often produce confident but incorrect answers in settings where abstention would be safer. Standard evaluation protocols, however, require a response a…
Entropy Alone is Insufficient for Safe Selective Prediction in LLMs
Edward Phillips, Fredrik K. Gustafsson, Sean Wu +2
Selective prediction systems can mitigate harms resulting from language model hallucinations by abstaining from answering in high-risk cases. Uncertainty quantification techniques…
AutoMedPrompt: A New Framework for Optimizing LLM Medical Prompts Using Textual Gradients
Sean Wu, Michael Koo, Fabien Scalzo +1
Large language models (LLMs) have demonstrated increasingly sophisticated performance in medical and other fields of knowledge. Traditional methods of creating specialist LLMs requ…