2 papers
cs.CV2025
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
Xinpeng Ding, Jianhua Han, Hang Xu +2
Recent efforts to use natural language for interpretable driving focus mainly on planning, neglecting perception tasks. In this paper, we address this gap by introducing ROLISP (Ri…
cs.CV2024
A Closer Look at the Explainability of Contrastive Language-Image Pre-training
Yi Li, Hualiang Wang, Yiqun Duan +2
Contrastive language-image pre-training (CLIP) is a powerful vision-language model that has shown great benefits for various tasks. However, we have identified some issues with its…