7 papers · 1 filter
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
Wenrui Zhou, Mohamed Hendy, Shu Yang +5
As video large language models (Video-LLMs) become increasingly integrated into real-world applications that demand grounded multimodal reasoning, ensuring their factual consistenc…
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
Zhipeng Yang, Shu Yang, Lijie Hu +1
Large language models (LLMs) trained with canonical tokenization exhibit surprising robustness to non-canonical inputs such as character-level tokenization, yet the mechanisms unde…
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
Hongji Li, Junchi yao, Manjiang Yu +4
Machine unlearning aims to erase requested data from trained models without full retraining. For Reasoning Multimodal Large Language Models (RMLLMs), this is uniquely challenging:…
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
Keyu Wang, Jin Li, Shu Yang +2
Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented th…
The Compositional Architecture of Regret in Large Language Models
Xiangxiang Cui, Shu Yang, Tianjin Huang +3
Regret in Large Language Models refers to their explicit regret expression when presented with evidence contradicting their previously generated misinformation. Studying the regret…
Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
Wenshuo Dong, Qingsong Yang, Shu Yang +5
Large Language Models (LLMs) trained on massive data capture rich information embedded in the training data. However, this also introduces the risk of privacy leakage, particularly…