activity
20242026
collaborators

7 papers

cs.CL2026

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

Yuxi Xia, Dennis Ulmer, Terra Blevins +3

Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the align…

cs.CL2026

Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis

Yuxi Xia, Kinga Stańczak, Benjamin Roth

AI-text detectors achieve high accuracy on in-domain benchmarks, but often struggle to generalize across different generation conditions such as unseen prompts, model families, or…

cs.LG2026

An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text

Loris Schoenegger, Yuxi Xia, Benjamin Roth

The increasing difficulty to distinguish language-model-generated from human-written text has led to the development of detectors of machine-generated text (MGT). However, in many…

cs.CL2026

Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs

Yuxi Xia, Loris Schoenegger, Benjamin Roth

Large language models (LLMs) can increase users' perceived trust by verbalizing confidence in their outputs. However, prior work has shown that LLMs are often overconfident, making…

cs.CL2025

Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles

Yuxi Xia, Pedro Henrique Luz de Araujo, Klim Zaporojets +1

Calibration, the alignment between model confidence and prediction accuracy, is critical for the reliable deployment of large language models (LLMs). Existing works neglect to meas…

cs.AI2025

Specification Overfitting in Artificial Intelligence

Benjamin Roth, Pedro Henrique Luz de Araujo, Yuxi Xia +2

Machine learning (ML) and artificial intelligence (AI) approaches are often criticized for their inherent bias and for their lack of control, accountability, and transparency. Cons…