1 citations · 1 across the 11 of their papers we have counts for
18 papers · 1 filter
FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams
Genying Li, Boda Lin, Jiachen Li +5
Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to…
MASS: Multiplayer World Models with Authoritative Shared State
Ziqi Cai, Siqi Yang, Yimu Wang +6
Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsisten…
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Kaiqi Liu, Yunyao Mao, Ziqi Cai +8
While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing met…
Video Generation Models Are Inherent Lighting Estimators
Ziqi Cai, Shuchen Weng, Kaiqi Liu +5
Recovering dynamic environment maps from a single in-the-wild video is crucial for photorealistic rendering, yet remains a challenge. Recent video generation models can produce pho…
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
Haojie Zheng, Yixin Yang, Siqi Yang +2
Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed…
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
Peixuan Zhang, Chang Zhou, Ziyuan Zhang +8
The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems. However, existing compilatio…