multimodal large language models 2data synthesis 1efficient video processing 1generalist models 1proactive reminder 1video benchmark 1video understanding 1visual impairment assistance 1visual question answering 1
From the 2 of 3 linked papers with an AI index.
3 papers
cs.CV2026
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal gro…
cs.CV2026
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…
cs.CV2026
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Yunfeng Liu, Yuandong Yang, Jiarui Han +5
The paper introduces VIABench, a video benchmark built from first‑person recordings by visually impaired users to evaluate multimodal large language models on tasks like proactive…