Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs
Yumeng Shi, Quanyu Long, Yin Wu +1
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token rep…
cs.CV2025
Causality Matters: How Temporal Information Emerges in Video Language Models
Yumeng Shi, Quanyu Long, Yin Wu +1
Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and…