collaborators

9 papers

cs.AI2025

Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment

Samuel Yeh, Sharon Li

Human feedback plays a pivotal role in aligning large language models (LLMs) with human preferences. However, such feedback is often noisy or inconsistent, which can degrade the qu…

cs.CL2025

Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models

Xuanming Zhang, Yuxuan Chen, Samuel Yeh +1

Large language models (LLMs) excel at complex reasoning but can still exhibit harmful behaviors. Current alignment strategies typically embed safety into model weights, making thes…

cs.CL2025

LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions

Yang Xu, Xuanming Zhang, Samuel Yeh +4

Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most eval…

cs.CL2025

LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals

Samuel Yeh, Sharon Li, Tanwi Mallick

Retrieval-Augmented Generation (RAG) aims to mitigate hallucinations in large language models (LLMs) by grounding responses in retrieved documents. Yet, RAG-based LLMs still halluc…

cs.CV2025

GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity

Seongheon Park, Sharon Li

Object hallucination in large vision-language models presents a significant challenge to their safe deployment in real-world applications. Recent works have proposed object-level h…

cs.AI2025

Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations

Jinyuan Luo, Zhen Fang, Yixuan Li +2

Hallucination remains a key obstacle to the reliable deployment of large language models (LLMs) in real-world question answering tasks. A widely adopted strategy to detect hallucin…