3 papers
cs.CV2026
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
Yi Liu, Jing Zhang, Di Wang +3
Multimodal large language models (MLLMs) suffer from pronounced hallucinations in remote sensing visual question-answering (RS-VQA), primarily caused by visual grounding failures i…
cs.CV2025
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
Zeyu Xu, Junkang Zhang, Qiang Wang +1
Vision-Language Models (VLMs) have enabled substantial progress in video understanding by leveraging cross-modal reasoning capabilities. However, their effectiveness is limited by…
cs.CV2025
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
Yi Liu, Xiao Xu, Zeyu Xu +10
Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computatio…