6 papers
When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge
Xin Sun, Di Wu, Yuchen Guo +4
LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidenc…
When LLM Rationales Become User-Facing: Effects on Trust Perception, Decision-Making, and Gaze Behaviors
Xin Sun, Ting Pan, Yajing Wang +5
Large language models (LLMs) increasingly show step-by-step reasoning rationales alongside their answers, turning reasoning from an internal model capability into a user-facing int…
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
Xin Sun, Di Wu, Sijing Qin +3
Large language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased…
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
Futa Waseda, Saku Sugawara, Isao Echizen
Defending pre-trained vision-language models (VLMs), such as CLIP, against adversarial attacks is crucial, as these models are widely used in diverse zero-shot tasks, including ima…
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
Yuchen Guo, Zhicheng Dou, Huy H. Nguyen +3
Content creation has dramatically progressed with the rapid advancement of large language models like ChatGPT and Claude. While this progress has greatly enhanced various aspects o…
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
Momoka Furuhashi, Hiroaki Funayama, Yuya Iwase +5
Short-reading comprehension questions help students understand text structure but lack effective feedback. Students struggle to identify and correct errors, while manual feedback c…