2 papers
cs.CV2025
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Yun Li, Zhe Liu, Yajing Kong +6
Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approac…
cs.CV2024
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
Minyi Zhao, Jie Wang, Zhaoyang Li +3
Recent studies have shown that Vision Language Large Models (VLLMs) may output content not relevant to the input images. This problem, called the hallucination phenomenon, undoubte…