3 papers
cs.CV2026
Geometry-Grounded Unified 3D Perception for Autonomous Driving
Longfei Xu, Xiaohui Wang, Zehao Huang +4
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-bas…
cs.AI2026
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
Zehao Deng, Tianjie Ju, Zheng Wu +5
Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the internal mechanisms governing how they encode and ground distinct visual concep…
cs.CV2025
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
Meng Luo, Shengqiong Wu, Liqiang Jing +12
Recent advancements in large video models (LVMs) have significantly enhance video understanding. However, these models continue to suffer from hallucinations, producing content tha…