collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

Yumeng Shi, Quanyu Long, Yin Wu +1

Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token rep…

cs.CV2026

Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention

Chengtao Lv, Yumeng Shi, Yushi Huang +3

Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention remains a primary bottleneck for eff…

cs.CV2025

Causality Matters: How Temporal Information Emerges in Video Language Models

Yumeng Shi, Quanyu Long, Yin Wu +1

Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and…

cs.CV2025

LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit

Chengtao Lv, Bilang Zhang, Yang Yong +7

Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequenc…

cs.CV2025

Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering

Yumeng Shi, Quanyu Long, Wenya Wang

Video question answering benefits from the rich information in videos, enabling various applications. However, the large volume of tokens generated from long videos presents challe…