collaborators

10 papers

cs.CV2026

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Jianzong Wu, Hao Lian, Jiongfan Yang +12

Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified framewo…

cs.LG2026

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

Qingyu Shi, Jinbin Bai, Zhuoran Zhao +7

Unified generation models aim to handle diverse tasks across modalities -- such as text generation, image generation, and vision-language reasoning -- within a single architecture…

cs.LG2026

Towards Customized Multimodal Role-Play

Chao Tang, Jianzong Wu, Qingyu Shi +5

Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while…

cs.CV2025

Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation

Jianzong Wu, Hao Lian, Dachao Hao +5

Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: D…

cs.CV2025

Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer

Qingyu Shi, Jianzong Wu, Jinbin Bai +4

The motion transfer task aims to transfer motion from a source video to newly generated videos, requiring the model to decouple motion from appearance. Previous diffusion-based met…

cs.CV2025

VMoBA: Mixture-of-Block Attention for Video Diffusion Models

Jianzong Wu, Liang Hou, Haotian Yang +5

The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high-resolution videos. Whi…