2 papers
cs.CV2026
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
Zhenguo Zhang, Haohan Zheng, Yishen Wang +6
The deployment of Vision-Language Models (VLMs) in safety-critical domains like autonomous driving (AD) is critically hindered by reliability failures, most notably object hallucin…
cs.CV2025
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
Haohan Zheng, Zhenguo Zhang
Large vision-language models (LVLMs) have demonstrated remarkable multimodal comprehension and reasoning capabilities, but they still suffer from severe object hallucination. Previ…