12 papers
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Kaixin Ding, Xi Chen, Minghong Cai +9
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllabi…
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4
Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this inter…
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
Junhao Cheng, Liang Hou, Tianxiong Zhong +4
The recent "Reasoning with Video" paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reasoning tasks. Although state-o…
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
Xinyu Mao, Yuhui Zeng, Xiaokun Liu +6
Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and sho…
Geometry-Instructed Video Editing
Chirui Chang, Xiaoyang Lyu, Yi-Hua Huang +7
Object-level geometric edits, including translating, rotating, scaling, duplicating, or removing an object, are routine operations in digital content creation (DCC) workflows, yet…
UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating
Jiehui Huang, Yuechen Zhang, Bin Xia +7
Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches…