works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.CV2026

NeMo: Needle in a Montage for Video-Language Understanding

Zi-Yuan Hu, Shuo Liang, Duo Zheng +10

The paper introduces the Needle in a Montage (NeMo) task and the NeMoBench benchmark to evaluate temporal understanding in video-language models, using an automated pipeline to gen…

cs.CV2026

VideoLatent: Video-Language Learning via Latent Self-Forcing

Zi-Yuan Hu, Zicong Tang, Shijia Huang +3

Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs).…

cs.CV2025

Rethinking Chain-of-Thought Reasoning for Videos

Yiwu Zhong, Zi-Yuan Hu, Yin Li +1

Chain-of-thought (CoT) reasoning has been highly successful in solving complex tasks in natural language processing, and recent multimodal large language models (MLLMs) have extend…

cs.CV2025

Fine-grained Spatiotemporal Grounding on Egocentric Videos

Shuo Liang, Yiwu Zhong, Zi-Yuan Hu +2

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the…

cs.CV2024

Enhancing Temporal Modeling of Video LLMs via Time Gating

Zi-Yuan Hu, Yiwu Zhong, Shijia Huang +2

Video Large Language Models (Video LLMs) have achieved impressive performance on video-and-language tasks, such as video question answering. However, most existing Video LLMs negle…

cs.CV2024

Beyond Embeddings: The Promise of Visual Table in Visual Reasoning

Yiwu Zhong, Zi-Yuan Hu, Michael R. Lyu +1

Visual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations. Des…