1 citations · 1 across the 5 of their papers we have counts for
6 papers · 1 filter
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
Yuxi Xia, Kinga Stańczak, Benjamin Roth
AI-text detectors achieve high accuracy on in-domain benchmarks, but often struggle to generalize across different generation conditions such as unseen prompts, model families, or…
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
Yuxi Xia, Loris Schoenegger, Benjamin Roth
Large language models (LLMs) can increase users' perceived trust by verbalizing confidence in their outputs. However, prior work has shown that LLMs are often overconfident, making…
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
Yuxi Xia, Dennis Ulmer, Terra Blevins +3
Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the align…
Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles
Yuxi Xia, Pedro Henrique Luz de Araujo, Klim Zaporojets +1
Calibration, the alignment between model confidence and prediction accuracy, is critical for the reliable deployment of large language models (LLMs). Existing works neglect to meas…
Black-box Model Ensembling for Textual and Visual Question Answering via Information Fusion
Yuxi Xia, Kilm Zaporojets, Benjamin Roth
A diverse range of large language models (LLMs), e.g., ChatGPT, and visual question answering (VQA) models, e.g., BLIP, have been developed for solving textual and visual question…
Exploring prompts to elicit memorization in masked language model-based named entity recognition
Yuxi Xia, Anastasiia Sedova, Pedro Henrique Luz de Araujo +3
Training data memorization in language models impacts model capability (generalization) and safety (privacy risk). This paper focuses on analyzing prompts' impact on detecting the…