9 papers
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
Samuel Yeh, Sharon Li
Human feedback plays a pivotal role in aligning large language models (LLMs) with human preferences. However, such feedback is often noisy or inconsistent, which can degrade the qu…
Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
Xuanming Zhang, Yuxuan Chen, Samuel Yeh +1
Large language models (LLMs) excel at complex reasoning but can still exhibit harmful behaviors. Current alignment strategies typically embed safety into model weights, making thes…
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
Yang Xu, Xuanming Zhang, Samuel Yeh +4
Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most eval…
LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
Samuel Yeh, Sharon Li, Tanwi Mallick
Retrieval-Augmented Generation (RAG) aims to mitigate hallucinations in large language models (LLMs) by grounding responses in retrieved documents. Yet, RAG-based LLMs still halluc…
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
Seongheon Park, Sharon Li
Object hallucination in large vision-language models presents a significant challenge to their safe deployment in real-world applications. Recent works have proposed object-level h…
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
Jinyuan Luo, Zhen Fang, Yixuan Li +2
Hallucination remains a key obstacle to the reliable deployment of large language models (LLMs) in real-world question answering tasks. A widely adopted strategy to detect hallucin…