4 papers · 1 filter
The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs
Yumeng Shi, Quanyu Long, Yin Wu +1
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token rep…
Causality Matters: How Temporal Information Emerges in Video Language Models
Yumeng Shi, Quanyu Long, Yin Wu +1
Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and…
LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
Chengtao Lv, Bilang Zhang, Yang Yong +7
Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequenc…
Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
Yumeng Shi, Quanyu Long, Wenya Wang
Video question answering benefits from the rich information in videos, enabling various applications. However, the large volume of tokens generated from long videos presents challe…