3 papers
cs.CV2026
Hallucination-aware intermediate representation edit in large vision-language models
Wei Suo, Hanzu Zhang, Lijun Zhang +3
Large Vision-Language Models have demonstrated exceptional performance in multimodal reasoning and complex scene understanding. However, these models still face significant halluci…
cs.CV2026
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
Ji Ma, Wei Suo, Peng Wang +1
Multimodal Chain-of-Thought (MCoT) models have demonstrated impressive capability in complex visual reasoning tasks. Unfortunately, recent studies reveal that they suffer from seve…
cs.CV2026
Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
Yizhou Jin, Yuezhu Feng, Jinjin Zhang +3
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to…