7 papers
Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?
Arnisa Fazla, Alberto Testoni, Ameen Abu-Hanna +2
Safe deployment of clinical vision-language models (VLMs) requires reliable uncertainty estimation (UE): a signal indicating when predictions should be trusted or escalated to a cl…
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
Benedetta Muscato, Beiduo Chen, Gizem Gezici +2
Human disagreement is ubiquitous and well-known in labeling. However, variation in explanations, captured through token-level human rationales, remains far less explored. At the sa…
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
Beiduo Chen, Pingjun Hong, Ziyun Zhang +3
Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisions. We study whether large…
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
Joris Baan, Wilker Aziz, Barbara Plank +1
Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling K responses yields a belief state - res…
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
Alberto Testoni, Barbara Plank, Raquel Fernández
Ambiguity resolution is key to effective communication. While humans effortlessly address ambiguity through conversational grounding strategies, the extent to which current languag…
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
Silvia Casola, Yang Janet Liu, Siyao Peng +3
Human language production exhibits remarkable richness and variation, reflecting diverse communication styles and intents. However, this variation is often overlooked in summarizat…