collaborators
Showing cs.CVShow all

37 papers · 1 filter

cs.CV2026

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Yifan Xu, Zihao Wang, Zhixiao Wang +6

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occ…

cs.CV2026

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

Xu Guo, Zhengxuan Wei, Xinghui Li +11

Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…

cs.CV2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Qixun Wang, Yang Shi, Letian Cheng +11

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with…

cs.CV2026

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Xinyu Liu, Shihao Li, Weihong Lin +10

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordi…

cs.CV2026

MemLearner: Learning to Query Context memory for Video World Models

Jiwen Yu, Jianxiong Gao, Jianhong Bai +7

Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world mode…

cs.CV2026

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

Junhao Cheng, Liang Hou, Tianxiong Zhong +4

The recent "Reasoning with Video" paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reasoning tasks. Although state-o…