2 citations · 3 across the 14 of their papers we have counts for
3 papers · 1 filter
From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents
Can Zhang, Baofeng Zhang, Xiaotian Han +5
Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discri…
Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
Wenti Yin, Xiaotian Han, Junyuan Shang +5
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing t…
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
Longbin Ji, Xiaoxiong Liu, Junyuan Shang +4
Recent advances in video generation have been dominated by diffusion and flow-matching models, which produce high-quality results but remain computationally intensive and difficult…