From the 1 of 7 linked papers with an AI index.
7 papers
NeMo: Needle in a Montage for Video-Language Understanding
Zi-Yuan Hu, Shuo Liang, Duo Zheng +10
The paper introduces the Needle in a Montage (NeMo) task and the NeMoBench benchmark to evaluate temporal understanding in video-language models, using an automated pipeline to gen…
VideoLatent: Video-Language Learning via Latent Self-Forcing
Zi-Yuan Hu, Zicong Tang, Shijia Huang +3
Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs).…
Rethinking Chain-of-Thought Reasoning for Videos
Yiwu Zhong, Zi-Yuan Hu, Yin Li +1
Chain-of-thought (CoT) reasoning has been highly successful in solving complex tasks in natural language processing, and recent multimodal large language models (MLLMs) have extend…
Fine-grained Spatiotemporal Grounding on Egocentric Videos
Shuo Liang, Yiwu Zhong, Zi-Yuan Hu +2
Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the…
Enhancing Temporal Modeling of Video LLMs via Time Gating
Zi-Yuan Hu, Yiwu Zhong, Shijia Huang +2
Video Large Language Models (Video LLMs) have achieved impressive performance on video-and-language tasks, such as video question answering. However, most existing Video LLMs negle…
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
Yiwu Zhong, Zi-Yuan Hu, Michael R. Lyu +1
Visual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations. Des…