37 papers · 1 filter
PercepCap: Video Captioner with Structured Spatio-Temporal Perception
Yifan Xu, Zihao Wang, Zhixiao Wang +6
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occ…
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Xu Guo, Zhengxuan Wei, Xinghui Li +11
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Qixun Wang, Yang Shi, Letian Cheng +11
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with…
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Xinyu Liu, Shihao Li, Weihong Lin +10
Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordi…
MemLearner: Learning to Query Context memory for Video World Models
Jiwen Yu, Jianxiong Gao, Jianhong Bai +7
Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world mode…
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
Junhao Cheng, Liang Hou, Tianxiong Zhong +4
The recent "Reasoning with Video" paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reasoning tasks. Although state-o…