2 papers
cs.CV2025
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
Jiankang Wang, Zhihan Zhang, Zhihang Liu +4
Multimodal large language models (MLLMs) have made remarkable progress in either temporal or spatial localization. However, they struggle to perform spatio-temporal video grounding…
cs.CV2024
Hallucination Mitigation Prompts Long-term Video Understanding
Yiwei Sun, Zhihang Liu, Chuanbin Liu +3
Recently, multimodal large language models have made significant advancements in video understanding tasks. However, their ability to understand unprocessed long videos is very lim…