Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson +5
Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study…
cs.CL2026
BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
Sean Wu, Fredrik K. Gustafsson, Edward Phillips +3
Large language models (LLMs) often produce confident but incorrect answers in settings where abstention would be safer. Standard evaluation protocols, however, require a response a…
cs.CL2026
Entropy Alone is Insufficient for Safe Selective Prediction in LLMs
Edward Phillips, Fredrik K. Gustafsson, Sean Wu +2
Selective prediction systems can mitigate harms resulting from language model hallucinations by abstaining from answering in high-risk cases. Uncertainty quantification techniques…