9 papers
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
Hongji Li, Junchi yao, Manjiang Yu +4
Machine unlearning aims to erase requested data from trained models without full retraining. For Reasoning Multimodal Large Language Models (RMLLMs), this is uniquely challenging:…
PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
Manjiang Yu, Hongji Li, Priyanka Singh +3
Reliable behavior control is central to deploying large language models (LLMs) on the web. Activation steering offers a tuning-free route to align attributes (e.g., truthfulness) t…
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
Zhuoran Zhang, Tengyue Wang, Xilin Gong +4
Multimodal large language models (MLLMs) must resolve conflicts when different modalities provide contradictory information, a process we term modality following. Prior work measur…
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
Keyu Wang, Jin Li, Shu Yang +2
Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented th…
The Compositional Architecture of Regret in Large Language Models
Xiangxiang Cui, Shu Yang, Tianjin Huang +3
Regret in Large Language Models refers to their explicit regret expression when presented with evidence contradicting their previously generated misinformation. Studying the regret…
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
Liangliang You, Junchi Yao, Shu Yang +3
While multimodal large language models excel at various tasks, they still suffer from hallucinations, which limit their reliability and scalability for broader domain applications.…