7 papers
Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation
Yanbo Ding, Yijia Fan, Caihua Shan +8
Diffusion Transformers (DiTs) have become the dominant paradigm for high-fidelity video generation, yet their ability to perform high-level semantic planning remains limited. While…
Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
Yanbo Ding, Zhizhi Guo, Quanyue Song +4
Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although several streaming audio-video gen…
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
Quanyue Song, Yishan He, Yanbo Ding +4
Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for…
MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
Xirui Hu, Yanbo Ding, Jiahao Wang +4
Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Exi…
MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation
Yanbo Ding, Xirui Hu, Zhizhi Guo +6
Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits…
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
Zhengrong Yue, Shaobin Zhuang, Kunchang Li +2
Despite the recent advancement in video stylization, most existing methods struggle to render any video with complex transitions, based on an open style description of user query.…