From the 1 of 5 linked papers with an AI index.
5 papers
NeMo: Needle in a Montage for Video-Language Understanding
Zi-Yuan Hu, Shuo Liang, Duo Zheng +10
The paper introduces the Needle in a Montage (NeMo) task and the NeMoBench benchmark to evaluate temporal understanding in video-language models, using an automated pipeline to gen…
VideoLatent: Video-Language Learning via Latent Self-Forcing
Zi-Yuan Hu, Zicong Tang, Shijia Huang +3
Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs).…
Rethinking Chain-of-Thought Reasoning for Videos
Yiwu Zhong, Zi-Yuan Hu, Yin Li +1
Chain-of-thought (CoT) reasoning has been highly successful in solving complex tasks in natural language processing, and recent multimodal large language models (MLLMs) have extend…
Fine-grained Spatiotemporal Grounding on Egocentric Videos
Shuo Liang, Yiwu Zhong, Zi-Yuan Hu +2
Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the…
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Yiwu Zhong, Zhuoming Liu, Yin Li +1
Large language models (LLMs) have enabled the creation of multi-modal LLMs that exhibit strong comprehension of visual data such as images and videos. However, these models usually…