6 papers
Generative Refinement Networks for Visual Synthesis
Jian Han, Jinlai Liu, Jiahuan Wang +2
While diffusion models dominate the field of visual generation, they are computationally inefficient, applying a uniform computational effort regardless of different complexity. In…
Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics
Jiahua Wang, Leqi Zheng, Jialong Wu +2
World models simulate environmental dynamics to enable embodied agents to plan and reason about future states. While real-world perception is inherently multimodal, existing approa…
SkyReels-V3 Technique Report
Debang Li, Zhengcong Fei, Tuanhui Li +19
Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReel…
MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
Ziang Luo, Kangan Qian, Jiahua Wang +13
Vision-Language Models(VLMs) have demonstrated significant potential for end-to-end autonomous driving, yet a substantial gap remains between their current capabilities and the rel…
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
Zhengcong Fei, Hao Jiang, Di Qiu +8
The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present…
SkyReels-A2: Compose Anything in Video Diffusion Transformers
Zhengcong Fei, Debang Li, Di Qiu +8
This paper presents SkyReels-A2, a controllable video generation framework capable of assembling arbitrary visual elements (e.g., characters, objects, backgrounds) into synthesized…