From the 1 of 17 linked papers with an AI index.
17 papers
MASS: Multiplayer World Models with Authoritative Shared State
Ziqi Cai, Siqi Yang, Yimu Wang +6
Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsisten…
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Kaiqi Liu, Yunyao Mao, Ziqi Cai +8
MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…
Video Generation Models Are Inherent Lighting Estimators
Ziqi Cai, Shuchen Weng, Kaiqi Liu +5
Recovering dynamic environment maps from a single in-the-wild video is crucial for photorealistic rendering, yet remains a challenge. Recent video generation models can produce pho…
MotionVLA: Vision-Language-Action Model for Humanoid Motion
Nonghai Zhang, Siyu Zhai, Yanjun Li +5
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods toke…
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
Haojie Zheng, Yixin Yang, Siqi Yang +2
Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed…
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
Haojie Zheng, Shuchen Weng, Jingqi Liu +3
Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, existing video editing methods larg…