11 papers
Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
Pingjun Hong, Benjamin Roth
Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support simulation in different ways.…
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
Beiduo Chen, Pingjun Hong, Ziyun Zhang +3
Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisions. We study whether large…
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
Yuxi Xia, Dennis Ulmer, Terra Blevins +3
Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the align…
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
Yuxi Xia, Kinga StaÅczak, Benjamin Roth
AI-text detectors achieve high accuracy on in-domain benchmarks, but often struggle to generalize across different generation conditions such as unseen prompts, model families, or…
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
Pedro Henrique Luz de Araujo, Michael A. Hedderich, Ali Modarressi +2
Persona-assigned large language models (LLMs) are used in domains such as education, healthcare, and sociodemographic simulation. Yet, they are typically evaluated only in short, s…
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
Loris Schoenegger, Yuxi Xia, Benjamin Roth
The increasing difficulty to distinguish language-model-generated from human-written text has led to the development of detectors of machine-generated text (MGT). However, in many…