works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

NeMo: Needle in a Montage for Video-Language Understanding

Zi-Yuan Hu, Shuo Liang, Duo Zheng +10

The paper introduces the Needle in a Montage (NeMo) task and the NeMoBench benchmark to evaluate temporal understanding in video-language models, using an automated pipeline to gen…

cs.CV2026

VideoLatent: Video-Language Learning via Latent Self-Forcing

Zi-Yuan Hu, Zicong Tang, Shijia Huang +3

Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs).…

cs.CV2025

Rethinking Chain-of-Thought Reasoning for Videos

Yiwu Zhong, Zi-Yuan Hu, Yin Li +1

Chain-of-thought (CoT) reasoning has been highly successful in solving complex tasks in natural language processing, and recent multimodal large language models (MLLMs) have extend…

cs.CV2025

Fine-grained Spatiotemporal Grounding on Egocentric Videos

Shuo Liang, Yiwu Zhong, Zi-Yuan Hu +2

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the…

cs.CV2025

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Yiwu Zhong, Zhuoming Liu, Yin Li +1

Large language models (LLMs) have enabled the creation of multi-modal LLMs that exhibit strong comprehension of visual data such as images and videos. However, these models usually…