From the 2 of 22 linked papers with an AI index.
22 papers
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal gro…
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Yunfeng Liu, Yuandong Yang, Jiarui Han +5
The paper introduces VIABench, a video benchmark built from first‑person recordings by visually impaired users to evaluate multimodal large language models on tasks like proactive…
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
Tianxiang Jiang, Sheng Xia, Yicheng Xu +5
While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical…
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Ming Xie, Zizheng Huang, Xudong Tan +6
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…
FreeRet: MLLMs as Training-Free Retrievers
Yuhan Zhu, Xiangyu Zeng, Chenting Wang +6
Multimodal large language models (MLLMs) are emerging as versatile foundations for mixed-modality retrieval. Yet, they often require heavy post-hoc training to convert them into co…